AirQA: A Comprehensive QA Dataset for AI Research with Instance-Level Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Tiancheng, Cao, Ruisheng, Zhang, Yuxin, Kang, Zhangyi, Wang, Zijian, Wang, Chenrun, Luo, Yijie, Zheng, Hang, Qian, Lirong, Chen, Lu, Yu, Kai |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RJUA-QA: A Comprehensive QA Dataset for Urology
von: Lyu, Shiwei, et al.
Veröffentlicht: (2023)
von: Lyu, Shiwei, et al.
Veröffentlicht: (2023)
MoleculeQA: A Dataset to Evaluate Factual Accuracy in Molecular Comprehension
von: Lu, Xingyu, et al.
Veröffentlicht: (2024)
von: Lu, Xingyu, et al.
Veröffentlicht: (2024)
ViQA-COVID: COVID-19 Machine Reading Comprehension Dataset for Vietnamese
von: Nguyen-Phung, Hai-Chung, et al.
Veröffentlicht: (2025)
von: Nguyen-Phung, Hai-Chung, et al.
Veröffentlicht: (2025)
JBE-QA: Japanese Bar Exam QA Dataset for Assessing Legal Domain Knowledge
von: Cao, Zhihan, et al.
Veröffentlicht: (2025)
von: Cao, Zhihan, et al.
Veröffentlicht: (2025)
AgriQA Dataset
von: Eldem, Ayşe
Veröffentlicht: (2026)
von: Eldem, Ayşe
Veröffentlicht: (2026)
Syn-QA2: Evaluating False Assumptions in Long-tail Questions with Synthetic QA Datasets
von: Daswani, Ashwin, et al.
Veröffentlicht: (2024)
von: Daswani, Ashwin, et al.
Veröffentlicht: (2024)
NeuSym-RAG: Hybrid Neural Symbolic Retrieval with Multiview Structuring for PDF Question Answering
von: Cao, Ruisheng, et al.
Veröffentlicht: (2025)
von: Cao, Ruisheng, et al.
Veröffentlicht: (2025)
ArabicaQA: A Comprehensive Dataset for Arabic Question Answering
von: Abdallah, Abdelrahman, et al.
Veröffentlicht: (2024)
von: Abdallah, Abdelrahman, et al.
Veröffentlicht: (2024)
FictionalQA: A Dataset for Studying Memorization and Knowledge Acquisition
von: Kirchenbauer, John, et al.
Veröffentlicht: (2025)
von: Kirchenbauer, John, et al.
Veröffentlicht: (2025)
BioGraphletQA: Knowledge-Anchored Generation of Complex QA Datasets
von: Jonker, Richard A. A., et al.
Veröffentlicht: (2026)
von: Jonker, Richard A. A., et al.
Veröffentlicht: (2026)
ChiMDQA: Towards Comprehensive Chinese Document QA with Fine-grained Evaluation
von: Gao, Jing, et al.
Veröffentlicht: (2025)
von: Gao, Jing, et al.
Veröffentlicht: (2025)
SEC-QA: A Systematic Evaluation Corpus for Financial QA
von: Lai, Viet Dac, et al.
Veröffentlicht: (2024)
von: Lai, Viet Dac, et al.
Veröffentlicht: (2024)
DeepSearchQA: Bridging the Comprehensiveness Gap for Deep Research Agents
von: Gupta, Nikita, et al.
Veröffentlicht: (2026)
von: Gupta, Nikita, et al.
Veröffentlicht: (2026)
BugBlitz-AI: An Intelligent QA Assistant
von: Yao, Yi, et al.
Veröffentlicht: (2024)
von: Yao, Yi, et al.
Veröffentlicht: (2024)
DocVideoQA: Towards Comprehensive Understanding of Document-Centric Videos through Question Answering
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
DisastQA: A Comprehensive Benchmark for Evaluating Question Answering in Disaster Management
von: Chen, Zhitong, et al.
Veröffentlicht: (2026)
von: Chen, Zhitong, et al.
Veröffentlicht: (2026)
LiteraryQA: Towards Effective Evaluation of Long-document Narrative QA
von: Bonomo, Tommaso, et al.
Veröffentlicht: (2025)
von: Bonomo, Tommaso, et al.
Veröffentlicht: (2025)
HumMusQA: A Human-written Music Understanding QA Benchmark Dataset
von: Weck, Benno, et al.
Veröffentlicht: (2026)
von: Weck, Benno, et al.
Veröffentlicht: (2026)
Liver Fibrosis Quantification and Analysis: The LiQA Dataset and Baseline Method
von: Liu, Yuanye, et al.
Veröffentlicht: (2025)
von: Liu, Yuanye, et al.
Veröffentlicht: (2025)
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
von: Zuo, Yuxin, et al.
Veröffentlicht: (2025)
von: Zuo, Yuxin, et al.
Veröffentlicht: (2025)
Compound-QA: A Benchmark for Evaluating LLMs on Compound Questions
von: Hou, Yutao, et al.
Veröffentlicht: (2024)
von: Hou, Yutao, et al.
Veröffentlicht: (2024)
FunQA: Towards Surprising Video Comprehension
von: Xie, Binzhu, et al.
Veröffentlicht: (2023)
von: Xie, Binzhu, et al.
Veröffentlicht: (2023)
ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models
von: Wang, Yueqian, et al.
Veröffentlicht: (2025)
von: Wang, Yueqian, et al.
Veröffentlicht: (2025)
Synthetic Data-Driven Prompt Tuning for Financial QA over Tables and Documents
von: Yu, Yaoning, et al.
Veröffentlicht: (2025)
von: Yu, Yaoning, et al.
Veröffentlicht: (2025)
Selective QA over Conflicting Multi-Source Personal Memory: A Diagnostic Testbed and Method Comparison
von: Yang, Tiancheng, et al.
Veröffentlicht: (2026)
von: Yang, Tiancheng, et al.
Veröffentlicht: (2026)
PolQA: Polish Question Answering Dataset
von: Rybak, Piotr, et al.
Veröffentlicht: (2022)
von: Rybak, Piotr, et al.
Veröffentlicht: (2022)
KET-QA: A Dataset for Knowledge Enhanced Table Question Answering
von: Hu, Mengkang, et al.
Veröffentlicht: (2024)
von: Hu, Mengkang, et al.
Veröffentlicht: (2024)
DeQA-Doc: Adapting DeQA-Score to Document Image Quality Assessment
von: Gao, Junjie, et al.
Veröffentlicht: (2025)
von: Gao, Junjie, et al.
Veröffentlicht: (2025)
ProMQA-Assembly: Multimodal Procedural QA Dataset on Assembly
von: Hasegawa, Kimihiro, et al.
Veröffentlicht: (2025)
von: Hasegawa, Kimihiro, et al.
Veröffentlicht: (2025)
IntentionQA: A Benchmark for Evaluating Purchase Intention Comprehension Abilities of Language Models in E-commerce
von: Ding, Wenxuan, et al.
Veröffentlicht: (2024)
von: Ding, Wenxuan, et al.
Veröffentlicht: (2024)
ECG-Expert-QA: A Benchmark for Evaluating Medical Large Language Models in Heart Disease Diagnosis
von: Wang, Xu, et al.
Veröffentlicht: (2025)
von: Wang, Xu, et al.
Veröffentlicht: (2025)
KGARevion: An AI Agent for Knowledge-Intensive Biomedical QA
von: Su, Xiaorui, et al.
Veröffentlicht: (2024)
von: Su, Xiaorui, et al.
Veröffentlicht: (2024)
GRS-QA -- Graph Reasoning-Structured Question Answering Dataset
von: Pahilajani, Anish, et al.
Veröffentlicht: (2024)
von: Pahilajani, Anish, et al.
Veröffentlicht: (2024)
Advancing 3D Scene Understanding with MV-ScanQA Multi-View Reasoning Evaluation and TripAlign Pre-training Dataset
von: Mo, Wentao, et al.
Veröffentlicht: (2025)
von: Mo, Wentao, et al.
Veröffentlicht: (2025)
VlogQA: Task, Dataset, and Baseline Models for Vietnamese Spoken-Based Machine Reading Comprehension
von: Ngo, Thinh Phuoc, et al.
Veröffentlicht: (2024)
von: Ngo, Thinh Phuoc, et al.
Veröffentlicht: (2024)
SustainableQA: A Comprehensive Question Answering Dataset for Corporate Sustainability and EU Taxonomy Reporting
von: Ali, Mohammed, et al.
Veröffentlicht: (2025)
von: Ali, Mohammed, et al.
Veröffentlicht: (2025)
SimpleDevQA: Benchmarking Large Language Models on Development Knowledge QA
von: Zhang, Jing, et al.
Veröffentlicht: (2025)
von: Zhang, Jing, et al.
Veröffentlicht: (2025)
RepoQA: Evaluating Long Context Code Understanding
von: Liu, Jiawei, et al.
Veröffentlicht: (2024)
von: Liu, Jiawei, et al.
Veröffentlicht: (2024)
FoQA: A Faroese Question-Answering Dataset
von: Simonsen, Annika, et al.
Veröffentlicht: (2025)
von: Simonsen, Annika, et al.
Veröffentlicht: (2025)
QA-LIGN: Aligning LLMs through Constitutionally Decomposed QA
von: Dineen, Jacob, et al.
Veröffentlicht: (2025)
von: Dineen, Jacob, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
RJUA-QA: A Comprehensive QA Dataset for Urology
von: Lyu, Shiwei, et al.
Veröffentlicht: (2023) -
MoleculeQA: A Dataset to Evaluate Factual Accuracy in Molecular Comprehension
von: Lu, Xingyu, et al.
Veröffentlicht: (2024) -
ViQA-COVID: COVID-19 Machine Reading Comprehension Dataset for Vietnamese
von: Nguyen-Phung, Hai-Chung, et al.
Veröffentlicht: (2025) -
JBE-QA: Japanese Bar Exam QA Dataset for Assessing Legal Domain Knowledge
von: Cao, Zhihan, et al.
Veröffentlicht: (2025) -
AgriQA Dataset
von: Eldem, Ayşe
Veröffentlicht: (2026)