Benchmarking Concept-Spilling Across Languages in LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Badanin, Ilia, Dzenhaliou, Daniil, Schlag, Imanol |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SFT-then-RL Outperforms Mixed-Policy Methods for LLM Reasoning
di: Limozin, Alexis, et al.
Pubblicazione: (2026)
di: Limozin, Alexis, et al.
Pubblicazione: (2026)
Spilled Energy in Large Language Models
di: Minut, Adrian Robert, et al.
Pubblicazione: (2026)
di: Minut, Adrian Robert, et al.
Pubblicazione: (2026)
Adapting Multilingual LLMs to Low-Resource Languages with Knowledge Graphs via Adapters
di: Gurgurov, Daniil, et al.
Pubblicazione: (2024)
di: Gurgurov, Daniil, et al.
Pubblicazione: (2024)
Persona-Augmented Benchmarking: Evaluating LLMs Across Diverse Writing Styles
di: Truong, Kimberly Le, et al.
Pubblicazione: (2025)
di: Truong, Kimberly Le, et al.
Pubblicazione: (2025)
LIBERTy: A Causal Framework for Benchmarking Concept-Based Explanations of LLMs with Structural Counterfactuals
di: Toker, Gilat, et al.
Pubblicazione: (2026)
di: Toker, Gilat, et al.
Pubblicazione: (2026)
Framing Political Bias in Multilingual LLMs Across Pakistani Languages
di: Nadeem, Afrozah, et al.
Pubblicazione: (2025)
di: Nadeem, Afrozah, et al.
Pubblicazione: (2025)
ConceptMath: A Bilingual Concept-wise Benchmark for Measuring Mathematical Reasoning of Large Language Models
di: Wu, Yanan, et al.
Pubblicazione: (2024)
di: Wu, Yanan, et al.
Pubblicazione: (2024)
Learning the Topic, Not the Language: How LLMs Classify Online Immigration Discourse Across Languages
di: Nasuto, Andrea, et al.
Pubblicazione: (2025)
di: Nasuto, Andrea, et al.
Pubblicazione: (2025)
Concept Attractors in LLMs and their Applications
di: Chytas, Sotirios Panagiotis, et al.
Pubblicazione: (2025)
di: Chytas, Sotirios Panagiotis, et al.
Pubblicazione: (2025)
<think> So let's replace this phrase with insult... </think> Lessons learned from generation of toxic texts with LLMs
di: Pletenev, Sergey, et al.
Pubblicazione: (2025)
di: Pletenev, Sergey, et al.
Pubblicazione: (2025)
Traveling Across Languages: Benchmarking Cross-Lingual Consistency in Multimodal LLMs
di: Wang, Hao, et al.
Pubblicazione: (2025)
di: Wang, Hao, et al.
Pubblicazione: (2025)
Beyond Specialization: Benchmarking LLMs for Transliteration of Indian Languages
di: Azam, Gulfarogh, et al.
Pubblicazione: (2025)
di: Azam, Gulfarogh, et al.
Pubblicazione: (2025)
Benchmarking LLMs for Mimicking Child-Caregiver Language in Interaction
di: Liu, Jing, et al.
Pubblicazione: (2024)
di: Liu, Jing, et al.
Pubblicazione: (2024)
NATURAL PLAN: Benchmarking LLMs on Natural Language Planning
di: Zheng, Huaixiu Steven, et al.
Pubblicazione: (2024)
di: Zheng, Huaixiu Steven, et al.
Pubblicazione: (2024)
POLAR-Bench: A Diagnostic Benchmark for Privacy-Utility Trade-offs in LLM Agents
di: Zheng, Qiaoyuan, et al.
Pubblicazione: (2026)
di: Zheng, Qiaoyuan, et al.
Pubblicazione: (2026)
Time Awareness in Large Language Models: Benchmarking Fact Recall Across Time
di: Herel, David, et al.
Pubblicazione: (2024)
di: Herel, David, et al.
Pubblicazione: (2024)
Cards Against LLMs: Benchmarking Humor Alignment in Large Language Models
di: Fettach, Yousra, et al.
Pubblicazione: (2026)
di: Fettach, Yousra, et al.
Pubblicazione: (2026)
General365: Benchmarking General Reasoning in Large Language Models Across Diverse and Challenging Tasks
di: Liu, Junlin, et al.
Pubblicazione: (2026)
di: Liu, Junlin, et al.
Pubblicazione: (2026)
Apertus: Democratizing Open and Compliant LLMs for Global Language Environments
di: Apertus, Project, et al.
Pubblicazione: (2025)
di: Apertus, Project, et al.
Pubblicazione: (2025)
Systematic Evaluation of Long-Context LLMs on Financial Concepts
di: Gupta, Lavanya, et al.
Pubblicazione: (2024)
di: Gupta, Lavanya, et al.
Pubblicazione: (2024)
Narrative Landscape: Mapping Narrative Dispositions Across LLMs
di: Jung, Donghoon, et al.
Pubblicazione: (2026)
di: Jung, Donghoon, et al.
Pubblicazione: (2026)
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels
di: Yan, Jianhao, et al.
Pubblicazione: (2024)
di: Yan, Jianhao, et al.
Pubblicazione: (2024)
Let's Play Across Cultures: A Large Multilingual, Multicultural Benchmark for Assessing Language Models' Understanding of Sports
di: Singh, Punit Kumar, et al.
Pubblicazione: (2025)
di: Singh, Punit Kumar, et al.
Pubblicazione: (2025)
Domain Knowledge-Enhanced LLMs for Fraud and Concept Drift Detection
di: Şenol, Ali, et al.
Pubblicazione: (2025)
di: Şenol, Ali, et al.
Pubblicazione: (2025)
LLM-BABYBENCH: Understanding and Evaluating Grounded Planning and Reasoning in LLMs
di: Choukrani, Omar, et al.
Pubblicazione: (2025)
di: Choukrani, Omar, et al.
Pubblicazione: (2025)
League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models
di: Guo, Qianhong, et al.
Pubblicazione: (2025)
di: Guo, Qianhong, et al.
Pubblicazione: (2025)
ConceptPsy:A Benchmark Suite with Conceptual Comprehensiveness in Psychology
di: Zhang, Junlei, et al.
Pubblicazione: (2023)
di: Zhang, Junlei, et al.
Pubblicazione: (2023)
Do Language Models Reason Across Languages?
di: Meng, Yan, et al.
Pubblicazione: (2026)
di: Meng, Yan, et al.
Pubblicazione: (2026)
Explanation Fairness in Large Language Models: An Empirical Analysis of Disparities in How LLMs Justify Decisions Across Demographic Groups
di: Veldanda, Gautam
Pubblicazione: (2026)
di: Veldanda, Gautam
Pubblicazione: (2026)
LLMs-in-the-Loop Part 2: Expert Small AI Models for Anonymization and De-identification of PHI Across Multiple Languages
di: Gunay, Murat, et al.
Pubblicazione: (2024)
di: Gunay, Murat, et al.
Pubblicazione: (2024)
Evaluating Modern Large Language Models on Low-Resource and Morphologically Rich Languages:A Cross-Lingual Benchmark Across Cantonese, Japanese, and Turkish
di: Xia, Chengxuan, et al.
Pubblicazione: (2025)
di: Xia, Chengxuan, et al.
Pubblicazione: (2025)
What is a Number, That a Large Language Model May Know It?
di: Marjieh, Raja, et al.
Pubblicazione: (2025)
di: Marjieh, Raja, et al.
Pubblicazione: (2025)
When Speculation Spills Secrets: Side Channels via Speculative Decoding In LLMs
di: Wei, Jiankun, et al.
Pubblicazione: (2024)
di: Wei, Jiankun, et al.
Pubblicazione: (2024)
CresOWLve: Benchmarking Creative Problem-Solving Over Real-World Knowledge
di: Ismayilzada, Mete, et al.
Pubblicazione: (2026)
di: Ismayilzada, Mete, et al.
Pubblicazione: (2026)
Soft Thinking: Unlocking the Reasoning Potential of LLMs in Continuous Concept Space
di: Zhang, Zhen, et al.
Pubblicazione: (2025)
di: Zhang, Zhen, et al.
Pubblicazione: (2025)
Facilitating large language model Russian adaptation with Learned Embedding Propagation
di: Tikhomirov, Mikhail, et al.
Pubblicazione: (2024)
di: Tikhomirov, Mikhail, et al.
Pubblicazione: (2024)
Predict the Next Word: Humans exhibit uncertainty in this task and language models _____
di: Ilia, Evgenia, et al.
Pubblicazione: (2024)
di: Ilia, Evgenia, et al.
Pubblicazione: (2024)
Can LLMs Correct Themselves? A Benchmark of Self-Correction in LLMs
di: Tie, Guiyao, et al.
Pubblicazione: (2025)
di: Tie, Guiyao, et al.
Pubblicazione: (2025)
Navigating the Concept Space of Language Models
di: Marcílio-Jr, Wilson E., et al.
Pubblicazione: (2026)
di: Marcílio-Jr, Wilson E., et al.
Pubblicazione: (2026)
Is Your LLM Really Mastering the Concept? A Multi-Agent Benchmark
di: Xu, Shuhang, et al.
Pubblicazione: (2025)
di: Xu, Shuhang, et al.
Pubblicazione: (2025)
Documenti analoghi
-
SFT-then-RL Outperforms Mixed-Policy Methods for LLM Reasoning
di: Limozin, Alexis, et al.
Pubblicazione: (2026) -
Spilled Energy in Large Language Models
di: Minut, Adrian Robert, et al.
Pubblicazione: (2026) -
Adapting Multilingual LLMs to Low-Resource Languages with Knowledge Graphs via Adapters
di: Gurgurov, Daniil, et al.
Pubblicazione: (2024) -
Persona-Augmented Benchmarking: Evaluating LLMs Across Diverse Writing Styles
di: Truong, Kimberly Le, et al.
Pubblicazione: (2025) -
LIBERTy: A Causal Framework for Benchmarking Concept-Based Explanations of LLMs with Structural Counterfactuals
di: Toker, Gilat, et al.
Pubblicazione: (2026)