MultiWikiQA: A Reading Comprehension Benchmark in 300+ Languages
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Smart, Dan Saattrup |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A multilingual hallucination benchmark: MultiWikiQHalluA
von: Thoresen, Freja, et al.
Veröffentlicht: (2026)
von: Thoresen, Freja, et al.
Veröffentlicht: (2026)
MultiZebraLogic: A Multilingual Logical Reasoning Benchmark
von: Bruun, Sofie Helene, et al.
Veröffentlicht: (2025)
von: Bruun, Sofie Helene, et al.
Veröffentlicht: (2025)
FoQA: A Faroese Question-Answering Dataset
von: Simonsen, Annika, et al.
Veröffentlicht: (2025)
von: Simonsen, Annika, et al.
Veröffentlicht: (2025)
JobResQA: A Benchmark for LLM Machine Reading Comprehension on Multilingual Résumés and JDs
von: Carrino, Casimiro Pio, et al.
Veröffentlicht: (2026)
von: Carrino, Casimiro Pio, et al.
Veröffentlicht: (2026)
WikiMixQA: A Multimodal Benchmark for Question Answering over Tables and Charts
von: Foroutan, Negar, et al.
Veröffentlicht: (2025)
von: Foroutan, Negar, et al.
Veröffentlicht: (2025)
Hotter and Colder: A New Approach to Annotating Sentiment, Emotions, and Bias in Icelandic Blog Comments
von: Friðriksdóttir, Steinunn Rut, et al.
Veröffentlicht: (2025)
von: Friðriksdóttir, Steinunn Rut, et al.
Veröffentlicht: (2025)
Encoder vs Decoder: Comparative Analysis of Encoder and Decoder Language Models on Multilingual NLU Tasks
von: Nielsen, Dan Saattrup, et al.
Veröffentlicht: (2024)
von: Nielsen, Dan Saattrup, et al.
Veröffentlicht: (2024)
ThReadMed-QA: A Multi-Turn Medical Dialogue Benchmark from Real Patient Questions
von: Munnangi, Monica, et al.
Veröffentlicht: (2026)
von: Munnangi, Monica, et al.
Veröffentlicht: (2026)
Medical Knowledge Graph QA for Drug-Drug Interaction Prediction based on Multi-hop Machine Reading Comprehension
von: Gao, Peng, et al.
Veröffentlicht: (2022)
von: Gao, Peng, et al.
Veröffentlicht: (2022)
IntentionQA: A Benchmark for Evaluating Purchase Intention Comprehension Abilities of Language Models in E-commerce
von: Ding, Wenxuan, et al.
Veröffentlicht: (2024)
von: Ding, Wenxuan, et al.
Veröffentlicht: (2024)
RJUA-QA: A Comprehensive QA Dataset for Urology
von: Lyu, Shiwei, et al.
Veröffentlicht: (2023)
von: Lyu, Shiwei, et al.
Veröffentlicht: (2023)
DisastQA: A Comprehensive Benchmark for Evaluating Question Answering in Disaster Management
von: Chen, Zhitong, et al.
Veröffentlicht: (2026)
von: Chen, Zhitong, et al.
Veröffentlicht: (2026)
MultiChartQA: Benchmarking Vision-Language Models on Multi-Chart Problems
von: Zhu, Zifeng, et al.
Veröffentlicht: (2024)
von: Zhu, Zifeng, et al.
Veröffentlicht: (2024)
ViQA-COVID: COVID-19 Machine Reading Comprehension Dataset for Vietnamese
von: Nguyen-Phung, Hai-Chung, et al.
Veröffentlicht: (2025)
von: Nguyen-Phung, Hai-Chung, et al.
Veröffentlicht: (2025)
VlogQA: Task, Dataset, and Baseline Models for Vietnamese Spoken-Based Machine Reading Comprehension
von: Ngo, Thinh Phuoc, et al.
Veröffentlicht: (2024)
von: Ngo, Thinh Phuoc, et al.
Veröffentlicht: (2024)
Beyond Facts: Benchmarking Distributional Reading Comprehension in Large Language Models
von: Guo, Pei-Fu, et al.
Veröffentlicht: (2026)
von: Guo, Pei-Fu, et al.
Veröffentlicht: (2026)
NaijaRC: A Multi-choice Reading Comprehension Dataset for Nigerian Languages
von: Aremu, Anuoluwapo, et al.
Veröffentlicht: (2023)
von: Aremu, Anuoluwapo, et al.
Veröffentlicht: (2023)
SportQA: A Benchmark for Sports Understanding in Large Language Models
von: Xia, Haotian, et al.
Veröffentlicht: (2024)
von: Xia, Haotian, et al.
Veröffentlicht: (2024)
MRCEval: A Comprehensive, Challenging and Accessible Machine Reading Comprehension Benchmark
von: Ma, Shengkun, et al.
Veröffentlicht: (2025)
von: Ma, Shengkun, et al.
Veröffentlicht: (2025)
M3SciQA: A Multi-Modal Multi-Document Scientific QA Benchmark for Evaluating Foundation Models
von: Li, Chuhan, et al.
Veröffentlicht: (2024)
von: Li, Chuhan, et al.
Veröffentlicht: (2024)
FanOutQA: A Multi-Hop, Multi-Document Question Answering Benchmark for Large Language Models
von: Zhu, Andrew, et al.
Veröffentlicht: (2024)
von: Zhu, Andrew, et al.
Veröffentlicht: (2024)
Wiki-Quantities and Wiki-Measurements: Datasets of Quantities and their Measurement Context from Wikipedia
von: Göpfert, Jan, et al.
Veröffentlicht: (2025)
von: Göpfert, Jan, et al.
Veröffentlicht: (2025)
ReasonTabQA: A Comprehensive Benchmark for Table Question Answering from Real World Industrial Scenarios
von: Pan, Changzai, et al.
Veröffentlicht: (2026)
von: Pan, Changzai, et al.
Veröffentlicht: (2026)
MedEthicsQA: A Comprehensive Question Answering Benchmark for Medical Ethics Evaluation of LLMs
von: Wei, Jianhui, et al.
Veröffentlicht: (2025)
von: Wei, Jianhui, et al.
Veröffentlicht: (2025)
HumMusQA: A Human-written Music Understanding QA Benchmark Dataset
von: Weck, Benno, et al.
Veröffentlicht: (2026)
von: Weck, Benno, et al.
Veröffentlicht: (2026)
AirQA: A Comprehensive QA Dataset for AI Research with Instance-Level Evaluation
von: Huang, Tiancheng, et al.
Veröffentlicht: (2025)
von: Huang, Tiancheng, et al.
Veröffentlicht: (2025)
M-QALM: A Benchmark to Assess Clinical Reading Comprehension and Knowledge Recall in Large Language Models via Question Answering
von: Subramanian, Anand, et al.
Veröffentlicht: (2024)
von: Subramanian, Anand, et al.
Veröffentlicht: (2024)
PaperHelper: Knowledge-Based LLM QA Paper Reading Assistant
von: Yin, Congrui, et al.
Veröffentlicht: (2025)
von: Yin, Congrui, et al.
Veröffentlicht: (2025)
ASL STEM Wiki: Dataset and Benchmark for Interpreting STEM Articles
von: Yin, Kayo, et al.
Veröffentlicht: (2024)
von: Yin, Kayo, et al.
Veröffentlicht: (2024)
MedConceptsQA: Open Source Medical Concepts QA Benchmark
von: Shoham, Ofir Ben, et al.
Veröffentlicht: (2024)
von: Shoham, Ofir Ben, et al.
Veröffentlicht: (2024)
MedExpQA: Multilingual Benchmarking of Large Language Models for Medical Question Answering
von: Alonso, Iñigo, et al.
Veröffentlicht: (2024)
von: Alonso, Iñigo, et al.
Veröffentlicht: (2024)
BaziQA-Benchmark: Evaluating Symbolic and Temporally Compositional Reasoning in Large Language Models
von: Chen, Jiangxi, et al.
Veröffentlicht: (2026)
von: Chen, Jiangxi, et al.
Veröffentlicht: (2026)
No Culture Left Behind: ArtELingo-28, a Benchmark of WikiArt with Captions in 28 Languages
von: Mohamed, Youssef, et al.
Veröffentlicht: (2024)
von: Mohamed, Youssef, et al.
Veröffentlicht: (2024)
GenomeQA: Benchmarking General Large Language Models for Genome Sequence Understanding
von: Long, Weicai, et al.
Veröffentlicht: (2026)
von: Long, Weicai, et al.
Veröffentlicht: (2026)
CodeReviewQA: The Code Review Comprehension Assessment for Large Language Models
von: Lin, Hong Yi, et al.
Veröffentlicht: (2025)
von: Lin, Hong Yi, et al.
Veröffentlicht: (2025)
DateLogicQA: Benchmarking Temporal Biases in Large Language Models
von: Bhatia, Gagan, et al.
Veröffentlicht: (2024)
von: Bhatia, Gagan, et al.
Veröffentlicht: (2024)
AfriMed-QA: A Pan-African, Multi-Specialty, Medical Question-Answering Benchmark Dataset
von: Olatunji, Tobi, et al.
Veröffentlicht: (2024)
von: Olatunji, Tobi, et al.
Veröffentlicht: (2024)
Compound-QA: A Benchmark for Evaluating LLMs on Compound Questions
von: Hou, Yutao, et al.
Veröffentlicht: (2024)
von: Hou, Yutao, et al.
Veröffentlicht: (2024)
MedFrameQA: A Multi-Image Medical VQA Benchmark for Clinical Reasoning
von: Yu, Suhao, et al.
Veröffentlicht: (2025)
von: Yu, Suhao, et al.
Veröffentlicht: (2025)
ChineseEcomQA: A Scalable E-commerce Concept Evaluation Benchmark for Large Language Models
von: Chen, Haibin, et al.
Veröffentlicht: (2025)
von: Chen, Haibin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A multilingual hallucination benchmark: MultiWikiQHalluA
von: Thoresen, Freja, et al.
Veröffentlicht: (2026) -
MultiZebraLogic: A Multilingual Logical Reasoning Benchmark
von: Bruun, Sofie Helene, et al.
Veröffentlicht: (2025) -
FoQA: A Faroese Question-Answering Dataset
von: Simonsen, Annika, et al.
Veröffentlicht: (2025) -
JobResQA: A Benchmark for LLM Machine Reading Comprehension on Multilingual Résumés and JDs
von: Carrino, Casimiro Pio, et al.
Veröffentlicht: (2026) -
WikiMixQA: A Multimodal Benchmark for Question Answering over Tables and Charts
von: Foroutan, Negar, et al.
Veröffentlicht: (2025)