Evaluating the Elementary Multilingual Capabilities of Large Language Models with MultiQ
Fuente:
arXiv
Saved in:
| Main Authors: | Holtermann, Carolin, Röttger, Paul, Dill, Timm, Lauscher, Anne |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SoS: Analysis of Surface over Semantics in Multilingual Text-To-Image Generation
by: Holtermann, Carolin, et al.
Published: (2026)
by: Holtermann, Carolin, et al.
Published: (2026)
TempViz: On the Evaluation of Temporal Knowledge in Text-to-Image Models
by: Holtermann, Carolin, et al.
Published: (2026)
by: Holtermann, Carolin, et al.
Published: (2026)
Around the World in 24 Hours: Probing LLM Knowledge of Time and Place
by: Holtermann, Carolin, et al.
Published: (2025)
by: Holtermann, Carolin, et al.
Published: (2025)
What the Weight?! A Unified Framework for Zero-Shot Knowledge Composition
by: Holtermann, Carolin, et al.
Published: (2024)
by: Holtermann, Carolin, et al.
Published: (2024)
ScaLearn: Simple and Highly Parameter-Efficient Task Transfer by Learning to Scale
by: Frohmann, Markus, et al.
Published: (2023)
by: Frohmann, Markus, et al.
Published: (2023)
Centurio: On Drivers of Multilingual Ability of Large Vision-Language Model
by: Geigle, Gregor, et al.
Published: (2025)
by: Geigle, Gregor, et al.
Published: (2025)
How Much Do LLMs Hallucinate across Languages? On Realistic Multilingual Estimation of LLM Hallucination
by: Islam, Saad Obaid ul, et al.
Published: (2025)
by: Islam, Saad Obaid ul, et al.
Published: (2025)
MultiQ&A: An Analysis in Measuring Robustness via Automated Crowdsourcing of Question Perturbations and Answers
by: Cho, Nicole, et al.
Published: (2025)
by: Cho, Nicole, et al.
Published: (2025)
SafetyPrompts: a Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety
by: Röttger, Paul, et al.
Published: (2024)
by: Röttger, Paul, et al.
Published: (2024)
Decoding Multilingual Moral Preferences: Unveiling LLM's Biases Through the Moral Machine Experiment
by: Vida, Karina, et al.
Published: (2024)
by: Vida, Karina, et al.
Published: (2024)
Towards Ethical Multi-Agent Systems of Large Language Models: A Mechanistic Interpretability Perspective
by: Lee, Jae Hee, et al.
Published: (2025)
by: Lee, Jae Hee, et al.
Published: (2025)
Large Language Models for Human-Machine Collaborative Particle Accelerator Tuning through Natural Language
by: Kaiser, Jan, et al.
Published: (2024)
by: Kaiser, Jan, et al.
Published: (2024)
GIMMICK -- Globally Inclusive Multimodal Multitask Cultural Knowledge Benchmarking
by: Schneider, Florian, et al.
Published: (2025)
by: Schneider, Florian, et al.
Published: (2025)
Large Language Models Discriminate Against Speakers of German Dialects
by: Bui, Minh Duc, et al.
Published: (2025)
by: Bui, Minh Duc, et al.
Published: (2025)
The Pluralistic Moral Gap: Understanding Judgment and Value Differences between Humans and Large Language Models
by: Russo, Giuseppe, et al.
Published: (2025)
by: Russo, Giuseppe, et al.
Published: (2025)
Just Go Parallel: Improving the Multilingual Capabilities of Large Language Models
by: Qorib, Muhammad Reza, et al.
Published: (2025)
by: Qorib, Muhammad Reza, et al.
Published: (2025)
Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models
by: Röttger, Paul, et al.
Published: (2024)
by: Röttger, Paul, et al.
Published: (2024)
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages
by: Han, Wenhan, et al.
Published: (2025)
by: Han, Wenhan, et al.
Published: (2025)
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
by: Röttger, Paul, et al.
Published: (2023)
by: Röttger, Paul, et al.
Published: (2023)
Evaluating Consistency and Reasoning Capabilities of Large Language Models
by: Saxena, Yash, et al.
Published: (2024)
by: Saxena, Yash, et al.
Published: (2024)
Improving Multilingual Capabilities with Cultural and Local Knowledge in Large Language Models While Enhancing Native Performance
by: Kadiyala, Ram Mohan Rao, et al.
Published: (2025)
by: Kadiyala, Ram Mohan Rao, et al.
Published: (2025)
Is It Good Data for Multilingual Instruction Tuning or Just Bad Multilingual Evaluation for Large Language Models?
by: Chen, Pinzhen, et al.
Published: (2024)
by: Chen, Pinzhen, et al.
Published: (2024)
ElectriQ: A Benchmark for Assessing the Response Capability of Large Language Models in Power Marketing
by: Wang, Jinzhi, et al.
Published: (2025)
by: Wang, Jinzhi, et al.
Published: (2025)
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues
by: Jang, Kyochul, et al.
Published: (2025)
by: Jang, Kyochul, et al.
Published: (2025)
MUG-Eval: A Proxy Evaluation Framework for Multilingual Generation Capabilities in Any Language
by: Song, Seyoung, et al.
Published: (2025)
by: Song, Seyoung, et al.
Published: (2025)
Evaluation Methodology for Large Language Models for Multilingual Document Question and Answer
by: Kahana, Adar, et al.
Published: (2024)
by: Kahana, Adar, et al.
Published: (2024)
Look at the Text: Instruction-Tuned Language Models are More Robust Multiple Choice Selectors than You Think
by: Wang, Xinpeng, et al.
Published: (2024)
by: Wang, Xinpeng, et al.
Published: (2024)
DialectalArabicMMLU: Benchmarking Dialectal Capabilities in Arabic and Multilingual Language Models
by: Altakrori, Malik H., et al.
Published: (2025)
by: Altakrori, Malik H., et al.
Published: (2025)
The Curious Case of Factual (Mis)Alignment between LLMs' Short- and Long-Form Answers
by: Islam, Saad Obaid ul, et al.
Published: (2025)
by: Islam, Saad Obaid ul, et al.
Published: (2025)
Sensitivity, Performance, Robustness: Deconstructing the Effect of Sociodemographic Prompting
by: Beck, Tilman, et al.
Published: (2023)
by: Beck, Tilman, et al.
Published: (2023)
Multilingual != Multicultural: Evaluating Gaps Between Multilingual Capabilities and Cultural Alignment in LLMs
by: Rystrøm, Jonathan, et al.
Published: (2025)
by: Rystrøm, Jonathan, et al.
Published: (2025)
SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors
by: Hu, Tiancheng, et al.
Published: (2025)
by: Hu, Tiancheng, et al.
Published: (2025)
FoundaBench: Evaluating Chinese Fundamental Knowledge Capabilities of Large Language Models
by: Li, Wei, et al.
Published: (2024)
by: Li, Wei, et al.
Published: (2024)
SQLBench: A Comprehensive Evaluation for Text-to-SQL Capabilities of Large Language Models
by: Zhang, Bin, et al.
Published: (2024)
by: Zhang, Bin, et al.
Published: (2024)
PCEval: A Benchmark for Evaluating Physical Computing Capabilities of Large Language Models
by: Song, Inpyo, et al.
Published: (2025)
by: Song, Inpyo, et al.
Published: (2025)
PerQ: Efficient Evaluation of Multilingual Text Personalization Quality
by: Macko, Dominik, et al.
Published: (2025)
by: Macko, Dominik, et al.
Published: (2025)
The Roles of English in Evaluating Multilingual Language Models
by: Poelman, Wessel, et al.
Published: (2024)
by: Poelman, Wessel, et al.
Published: (2024)
Multilingual Collaborative Defense for Large Language Models
by: Li, Hongliang, et al.
Published: (2025)
by: Li, Hongliang, et al.
Published: (2025)
MoZIP: A Multilingual Benchmark to Evaluate Large Language Models in Intellectual Property
by: Ni, Shiwen, et al.
Published: (2024)
by: Ni, Shiwen, et al.
Published: (2024)
The Heap: A Contamination-Free Multilingual Code Dataset for Evaluating Large Language Models
by: Katzy, Jonathan, et al.
Published: (2025)
by: Katzy, Jonathan, et al.
Published: (2025)
Similar Items
-
SoS: Analysis of Surface over Semantics in Multilingual Text-To-Image Generation
by: Holtermann, Carolin, et al.
Published: (2026) -
TempViz: On the Evaluation of Temporal Knowledge in Text-to-Image Models
by: Holtermann, Carolin, et al.
Published: (2026) -
Around the World in 24 Hours: Probing LLM Knowledge of Time and Place
by: Holtermann, Carolin, et al.
Published: (2025) -
What the Weight?! A Unified Framework for Zero-Shot Knowledge Composition
by: Holtermann, Carolin, et al.
Published: (2024) -
ScaLearn: Simple and Highly Parameter-Efficient Task Transfer by Learning to Scale
by: Frohmann, Markus, et al.
Published: (2023)