MultiLoKo: a multilingual local knowledge benchmark for LLMs spanning 31 languages
Fuente:
arXiv
Saved in:
| Main Authors: | Hupkes, Dieuwke, Bogoychev, Nikolay |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Form(s) to Meaning: Probing the Semantic Depths of Language Models Using Multisense Consistency
by: Ohmer, Xenia, et al.
Published: (2024)
by: Ohmer, Xenia, et al.
Published: (2024)
Interpretability of Language Models via Task Spaces
by: Weber, Lucas, et al.
Published: (2024)
by: Weber, Lucas, et al.
Published: (2024)
Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
by: Thakur, Aman Singh, et al.
Published: (2024)
by: Thakur, Aman Singh, et al.
Published: (2024)
Lost in Inference: Rediscovering the Role of Natural Language Inference for Large Language Models
by: Madaan, Lovish, et al.
Published: (2024)
by: Madaan, Lovish, et al.
Published: (2024)
The Ups and Downs of Large Language Model Inference with Vocabulary Trimming by Language Heuristics
by: Bogoychev, Nikolay, et al.
Published: (2023)
by: Bogoychev, Nikolay, et al.
Published: (2023)
Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?
by: Singh, Aaditya K., et al.
Published: (2024)
by: Singh, Aaditya K., et al.
Published: (2024)
Compute Optimal Scaling of Skills: Knowledge vs Reasoning
by: Roberts, Nicholas, et al.
Published: (2025)
by: Roberts, Nicholas, et al.
Published: (2025)
A multilingual hallucination benchmark: MultiWikiQHalluA
by: Thoresen, Freja, et al.
Published: (2026)
by: Thoresen, Freja, et al.
Published: (2026)
Halluverse-M^3: A multitask multilingual benchmark for hallucination in LLMs
by: Abdaljalil, Samir, et al.
Published: (2026)
by: Abdaljalil, Samir, et al.
Published: (2026)
Empirical study of pretrained multilingual language models for zero-shot cross-lingual knowledge transfer in generation
by: Chirkova, Nadezhda, et al.
Published: (2023)
by: Chirkova, Nadezhda, et al.
Published: (2023)
Monolingual or Multilingual Instruction Tuning: Which Makes a Better Alpaca
by: Chen, Pinzhen, et al.
Published: (2023)
by: Chen, Pinzhen, et al.
Published: (2023)
Effective vocabulary expanding of multilingual language models for extremely low-resource languages
by: Zheng, Jianyu
Published: (2026)
by: Zheng, Jianyu
Published: (2026)
Clinical named entity recognition in the Portuguese language: a benchmark of modern BERT models and LLMs
by: de Almeida, Vinicius Anjos, et al.
Published: (2026)
by: de Almeida, Vinicius Anjos, et al.
Published: (2026)
EuroGEST: Investigating gender stereotypes in multilingual language models
by: Rowe, Jacqueline, et al.
Published: (2025)
by: Rowe, Jacqueline, et al.
Published: (2025)
SRS-Stories: Vocabulary-constrained multilingual story generation for language learning
by: Kamzela, Wiktor, et al.
Published: (2025)
by: Kamzela, Wiktor, et al.
Published: (2025)
Understanding the effects of language-specific class imbalance in multilingual fine-tuning
by: Jung, Vincent, et al.
Published: (2024)
by: Jung, Vincent, et al.
Published: (2024)
Understanding the role of FFNs in driving multilingual behaviour in LLMs
by: Bhattacharya, Sunit, et al.
Published: (2024)
by: Bhattacharya, Sunit, et al.
Published: (2024)
Linguini: A benchmark for language-agnostic linguistic reasoning
by: Sánchez, Eduardo, et al.
Published: (2024)
by: Sánchez, Eduardo, et al.
Published: (2024)
Information availability in different languages and various technological constraints related to multilinguism on the Internet
by: Khosla, Sonal, et al.
Published: (2025)
by: Khosla, Sonal, et al.
Published: (2025)
Artificial intelligence language technologies in multilingual healthcare: Grand challenges ahead
by: Briva-Iglesias, Vicent
Published: (2026)
by: Briva-Iglesias, Vicent
Published: (2026)
KoSimpleQA: A Korean Factuality Benchmark with an Analysis of Reasoning LLMs
by: Ko, Donghyeon, et al.
Published: (2025)
by: Ko, Donghyeon, et al.
Published: (2025)
One ruler to measure them all: Benchmarking multilingual long-context language models
by: Kim, Yekyung, et al.
Published: (2025)
by: Kim, Yekyung, et al.
Published: (2025)
Scalable multilingual PII annotation for responsible AI in LLMs
by: Meena, Bharti, et al.
Published: (2025)
by: Meena, Bharti, et al.
Published: (2025)
GRASP: A novel benchmark for evaluating language GRounding And Situated Physics understanding in multimodal language models
by: Jassim, Serwan, et al.
Published: (2023)
by: Jassim, Serwan, et al.
Published: (2023)
\textsc{CantoNLU}: A benchmark for Cantonese natural language understanding
by: Min, Junghyun, et al.
Published: (2025)
by: Min, Junghyun, et al.
Published: (2025)
WorldMedQA-V: a multilingual, multimodal medical examination dataset for multimodal language models evaluation
by: Matos, João, et al.
Published: (2024)
by: Matos, João, et al.
Published: (2024)
KoBLEX: Open Legal Question Answering with Multi-hop Reasoning
by: Lee, Jihyung, et al.
Published: (2025)
by: Lee, Jihyung, et al.
Published: (2025)
Fine-tuning multilingual language models in Twitter/X sentiment analysis: a study on Eastern-European V4 languages
by: Filip, Tomáš, et al.
Published: (2024)
by: Filip, Tomáš, et al.
Published: (2024)
An Expert-grounded benchmark of General Purpose LLMs in LCA
by: Donaldson, Artur, et al.
Published: (2025)
by: Donaldson, Artur, et al.
Published: (2025)
Anchor function: a type of benchmark functions for studying language models
by: Zhang, Zhongwang, et al.
Published: (2024)
by: Zhang, Zhongwang, et al.
Published: (2024)
Making Sense of Korean Sentences: A Comprehensive Evaluation of LLMs through KoSEnd Dataset
by: Yu, Seunguk, et al.
Published: (2025)
by: Yu, Seunguk, et al.
Published: (2025)
A benchmark dataset for evaluating Syndrome Differentiation and Treatment in large language models
by: Li, Kunning, et al.
Published: (2025)
by: Li, Kunning, et al.
Published: (2025)
AIDBench: A benchmark for evaluating the authorship identification capability of large language models
by: Wen, Zichen, et al.
Published: (2024)
by: Wen, Zichen, et al.
Published: (2024)
DevBench: A multimodal developmental benchmark for language learning
by: Tan, Alvin Wei Ming, et al.
Published: (2024)
by: Tan, Alvin Wei Ming, et al.
Published: (2024)
Investigating the interaction of linguistic and mathematical reasoning in language models using multilingual number puzzles
by: Bhattacharya, Antara Raaghavi, et al.
Published: (2025)
by: Bhattacharya, Antara Raaghavi, et al.
Published: (2025)
MultiCaption: Detecting disinformation using multilingual visual claims
by: Frade, Rafael Martins, et al.
Published: (2026)
by: Frade, Rafael Martins, et al.
Published: (2026)
Polish-English medical knowledge transfer: A new benchmark and results
by: Grzybowski, Łukasz, et al.
Published: (2024)
by: Grzybowski, Łukasz, et al.
Published: (2024)
NLPre: a revised approach towards language-centric benchmarking of Natural Language Preprocessing systems
by: Wiącek, Martyna, et al.
Published: (2024)
by: Wiącek, Martyna, et al.
Published: (2024)
The Lucie-7B LLM and the Lucie Training Dataset: Open resources for multilingual language generation
by: Gouvert, Olivier, et al.
Published: (2025)
by: Gouvert, Olivier, et al.
Published: (2025)
Does language matter for spoken word classification? A multilingual generative meta-learning approach
by: Ziki, Batsirayi Mupamhi, et al.
Published: (2026)
by: Ziki, Batsirayi Mupamhi, et al.
Published: (2026)
Similar Items
-
From Form(s) to Meaning: Probing the Semantic Depths of Language Models Using Multisense Consistency
by: Ohmer, Xenia, et al.
Published: (2024) -
Interpretability of Language Models via Task Spaces
by: Weber, Lucas, et al.
Published: (2024) -
Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
by: Thakur, Aman Singh, et al.
Published: (2024) -
Lost in Inference: Rediscovering the Role of Natural Language Inference for Large Language Models
by: Madaan, Lovish, et al.
Published: (2024) -
The Ups and Downs of Large Language Model Inference with Vocabulary Trimming by Language Heuristics
by: Bogoychev, Nikolay, et al.
Published: (2023)