RespondeoQA: a Benchmark for Bilingual Latin-English Question Answering
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hudspeth, Marisa, Burns, Patrick J., O'Connor, Brendan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Contextual morphologically-guided tokenization for Latin encoder models
von: Hudspeth, Marisa, et al.
Veröffentlicht: (2025)
von: Hudspeth, Marisa, et al.
Veröffentlicht: (2025)
Latin Treebanks in Review: An Evaluation of Morphological Tagging Across Time
von: Hudspeth, Marisa, et al.
Veröffentlicht: (2024)
von: Hudspeth, Marisa, et al.
Veröffentlicht: (2024)
Evaluating Morphological Alignment of Tokenizers in 70 Languages
von: Arnett, Catherine, et al.
Veröffentlicht: (2025)
von: Arnett, Catherine, et al.
Veröffentlicht: (2025)
BEnQA: A Question Answering and Reasoning Benchmark for Bengali and English
von: Shafayat, Sheikh, et al.
Veröffentlicht: (2024)
von: Shafayat, Sheikh, et al.
Veröffentlicht: (2024)
A Monte Carlo Language Model Pipeline for Zero-Shot Sociopolitical Event Extraction
von: Cai, Erica, et al.
Veröffentlicht: (2023)
von: Cai, Erica, et al.
Veröffentlicht: (2023)
Coordinates from Context: Using LLMs to Ground Complex Location References
von: Masis, Tessa, et al.
Veröffentlicht: (2025)
von: Masis, Tessa, et al.
Veröffentlicht: (2025)
Where on Earth Do Users Say They Are?: Geo-Entity Linking for Noisy Multilingual User Input
von: Masis, Tessa, et al.
Veröffentlicht: (2024)
von: Masis, Tessa, et al.
Veröffentlicht: (2024)
DashboardQA: Benchmarking Multimodal Agents for Question Answering on Interactive Dashboards
von: Kartha, Aaryaman, et al.
Veröffentlicht: (2025)
von: Kartha, Aaryaman, et al.
Veröffentlicht: (2025)
Understanding the Effect of Knowledge Graph Extraction Error on Downstream Graph Analyses: A Case Study on Affiliation Graphs
von: Cai, Erica, et al.
Veröffentlicht: (2025)
von: Cai, Erica, et al.
Veröffentlicht: (2025)
DisastQA: A Comprehensive Benchmark for Evaluating Question Answering in Disaster Management
von: Chen, Zhitong, et al.
Veröffentlicht: (2026)
von: Chen, Zhitong, et al.
Veröffentlicht: (2026)
AmharicStoryQA: A Multicultural Story Question Answering Benchmark in Amharic
von: Azime, Israel Abebe, et al.
Veröffentlicht: (2026)
von: Azime, Israel Abebe, et al.
Veröffentlicht: (2026)
NovelQA: Benchmarking Question Answering on Documents Exceeding 200K Tokens
von: Wang, Cunxiang, et al.
Veröffentlicht: (2024)
von: Wang, Cunxiang, et al.
Veröffentlicht: (2024)
CondAmbigQA: A Benchmark and Dataset for Conditional Ambiguous Question Answering
von: Li, Zongxi, et al.
Veröffentlicht: (2025)
von: Li, Zongxi, et al.
Veröffentlicht: (2025)
ASTRA-QA: A Benchmark for Abstract Question Answering over Documents
von: Wang, Shu, et al.
Veröffentlicht: (2026)
von: Wang, Shu, et al.
Veröffentlicht: (2026)
SensorQA: A Question Answering Benchmark for Daily-Life Monitoring
von: Reichman, Benjamin, et al.
Veröffentlicht: (2025)
von: Reichman, Benjamin, et al.
Veröffentlicht: (2025)
MedExQA: Medical Question Answering Benchmark with Multiple Explanations
von: Kim, Yunsoo, et al.
Veröffentlicht: (2024)
von: Kim, Yunsoo, et al.
Veröffentlicht: (2024)
PolQA: Polish Question Answering Dataset
von: Rybak, Piotr, et al.
Veröffentlicht: (2022)
von: Rybak, Piotr, et al.
Veröffentlicht: (2022)
MedExpQA: Multilingual Benchmarking of Large Language Models for Medical Question Answering
von: Alonso, Iñigo, et al.
Veröffentlicht: (2024)
von: Alonso, Iñigo, et al.
Veröffentlicht: (2024)
FSLI: An Interpretable Formal Semantic System for One-Dimensional Ordering Inference
von: Alkhairy, Maha, et al.
Veröffentlicht: (2025)
von: Alkhairy, Maha, et al.
Veröffentlicht: (2025)
DebateQA: Evaluating Question Answering on Debatable Knowledge
von: Xu, Rongwu, et al.
Veröffentlicht: (2024)
von: Xu, Rongwu, et al.
Veröffentlicht: (2024)
MizanQA: Benchmarking Large Language Models on Moroccan Legal Question Answering
von: Bahaj, Adil, et al.
Veröffentlicht: (2025)
von: Bahaj, Adil, et al.
Veröffentlicht: (2025)
LaMP-QA: A Benchmark for Personalized Long-form Question Answering
von: Salemi, Alireza, et al.
Veröffentlicht: (2025)
von: Salemi, Alireza, et al.
Veröffentlicht: (2025)
M2QA: Multi-domain Multilingual Question Answering
von: Engländer, Leon, et al.
Veröffentlicht: (2024)
von: Engländer, Leon, et al.
Veröffentlicht: (2024)
GRS-QA -- Graph Reasoning-Structured Question Answering Dataset
von: Pahilajani, Anish, et al.
Veröffentlicht: (2024)
von: Pahilajani, Anish, et al.
Veröffentlicht: (2024)
MobQA: A Benchmark Dataset for Semantic Understanding of Human Mobility Data through Question Answering
von: Asano, Hikaru, et al.
Veröffentlicht: (2025)
von: Asano, Hikaru, et al.
Veröffentlicht: (2025)
AfriMed-QA: A Pan-African, Multi-Specialty, Medical Question-Answering Benchmark Dataset
von: Olatunji, Tobi, et al.
Veröffentlicht: (2024)
von: Olatunji, Tobi, et al.
Veröffentlicht: (2024)
FoQA: A Faroese Question-Answering Dataset
von: Simonsen, Annika, et al.
Veröffentlicht: (2025)
von: Simonsen, Annika, et al.
Veröffentlicht: (2025)
Evaluating Monolingual and Multilingual Large Language Models for Greek Question Answering: The DemosQA Benchmark
von: Mastrokostas, Charalampos, et al.
Veröffentlicht: (2026)
von: Mastrokostas, Charalampos, et al.
Veröffentlicht: (2026)
RepLiQA: A Question-Answering Dataset for Benchmarking LLMs on Unseen Reference Content
von: Monteiro, Joao, et al.
Veröffentlicht: (2024)
von: Monteiro, Joao, et al.
Veröffentlicht: (2024)
MedEthicsQA: A Comprehensive Question Answering Benchmark for Medical Ethics Evaluation of LLMs
von: Wei, Jianhui, et al.
Veröffentlicht: (2025)
von: Wei, Jianhui, et al.
Veröffentlicht: (2025)
ReasonTabQA: A Comprehensive Benchmark for Table Question Answering from Real World Industrial Scenarios
von: Pan, Changzai, et al.
Veröffentlicht: (2026)
von: Pan, Changzai, et al.
Veröffentlicht: (2026)
ReCoQA: A Benchmark for Tool-Augmented and Multi-Step Reasoning in Real Estate Question and Answering
von: Zhang, Yindong, et al.
Veröffentlicht: (2026)
von: Zhang, Yindong, et al.
Veröffentlicht: (2026)
EconLogicQA: A Question-Answering Benchmark for Evaluating Large Language Models in Economic Sequential Reasoning
von: Quan, Yinzhu, et al.
Veröffentlicht: (2024)
von: Quan, Yinzhu, et al.
Veröffentlicht: (2024)
WikiMixQA: A Multimodal Benchmark for Question Answering over Tables and Charts
von: Foroutan, Negar, et al.
Veröffentlicht: (2025)
von: Foroutan, Negar, et al.
Veröffentlicht: (2025)
HistoryBankQA: Multilingual Temporal Question Answering on Historical Events
von: Mandal, Biswadip, et al.
Veröffentlicht: (2025)
von: Mandal, Biswadip, et al.
Veröffentlicht: (2025)
NeoQA: Evidence-based Question Answering with Generated News Events
von: Glockner, Max, et al.
Veröffentlicht: (2025)
von: Glockner, Max, et al.
Veröffentlicht: (2025)
KET-QA: A Dataset for Knowledge Enhanced Table Question Answering
von: Hu, Mengkang, et al.
Veröffentlicht: (2024)
von: Hu, Mengkang, et al.
Veröffentlicht: (2024)
Question-Answering (QA) Model for a Personalized Learning Assistant for Arabic Language
von: Sammoudi, Mohammad, et al.
Veröffentlicht: (2024)
von: Sammoudi, Mohammad, et al.
Veröffentlicht: (2024)
Answering Questions in Stages: Prompt Chaining for Contract QA
von: Roegiest, Adam, et al.
Veröffentlicht: (2024)
von: Roegiest, Adam, et al.
Veröffentlicht: (2024)
ArabicaQA: A Comprehensive Dataset for Arabic Question Answering
von: Abdallah, Abdelrahman, et al.
Veröffentlicht: (2024)
von: Abdallah, Abdelrahman, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Contextual morphologically-guided tokenization for Latin encoder models
von: Hudspeth, Marisa, et al.
Veröffentlicht: (2025) -
Latin Treebanks in Review: An Evaluation of Morphological Tagging Across Time
von: Hudspeth, Marisa, et al.
Veröffentlicht: (2024) -
Evaluating Morphological Alignment of Tokenizers in 70 Languages
von: Arnett, Catherine, et al.
Veröffentlicht: (2025) -
BEnQA: A Question Answering and Reasoning Benchmark for Bengali and English
von: Shafayat, Sheikh, et al.
Veröffentlicht: (2024) -
A Monte Carlo Language Model Pipeline for Zero-Shot Sociopolitical Event Extraction
von: Cai, Erica, et al.
Veröffentlicht: (2023)