AirQA: A Comprehensive QA Dataset for AI Research with Instance-Level Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Tiancheng, Cao, Ruisheng, Zhang, Yuxin, Kang, Zhangyi, Wang, Zijian, Wang, Chenrun, Luo, Yijie, Zheng, Hang, Qian, Lirong, Chen, Lu, Yu, Kai |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RJUA-QA: A Comprehensive QA Dataset for Urology
by: Lyu, Shiwei, et al.
Published: (2023)
by: Lyu, Shiwei, et al.
Published: (2023)
MoleculeQA: A Dataset to Evaluate Factual Accuracy in Molecular Comprehension
by: Lu, Xingyu, et al.
Published: (2024)
by: Lu, Xingyu, et al.
Published: (2024)
ViQA-COVID: COVID-19 Machine Reading Comprehension Dataset for Vietnamese
by: Nguyen-Phung, Hai-Chung, et al.
Published: (2025)
by: Nguyen-Phung, Hai-Chung, et al.
Published: (2025)
JBE-QA: Japanese Bar Exam QA Dataset for Assessing Legal Domain Knowledge
by: Cao, Zhihan, et al.
Published: (2025)
by: Cao, Zhihan, et al.
Published: (2025)
AgriQA Dataset
by: Eldem, Ayşe
Published: (2026)
by: Eldem, Ayşe
Published: (2026)
Syn-QA2: Evaluating False Assumptions in Long-tail Questions with Synthetic QA Datasets
by: Daswani, Ashwin, et al.
Published: (2024)
by: Daswani, Ashwin, et al.
Published: (2024)
NeuSym-RAG: Hybrid Neural Symbolic Retrieval with Multiview Structuring for PDF Question Answering
by: Cao, Ruisheng, et al.
Published: (2025)
by: Cao, Ruisheng, et al.
Published: (2025)
ArabicaQA: A Comprehensive Dataset for Arabic Question Answering
by: Abdallah, Abdelrahman, et al.
Published: (2024)
by: Abdallah, Abdelrahman, et al.
Published: (2024)
FictionalQA: A Dataset for Studying Memorization and Knowledge Acquisition
by: Kirchenbauer, John, et al.
Published: (2025)
by: Kirchenbauer, John, et al.
Published: (2025)
BioGraphletQA: Knowledge-Anchored Generation of Complex QA Datasets
by: Jonker, Richard A. A., et al.
Published: (2026)
by: Jonker, Richard A. A., et al.
Published: (2026)
ChiMDQA: Towards Comprehensive Chinese Document QA with Fine-grained Evaluation
by: Gao, Jing, et al.
Published: (2025)
by: Gao, Jing, et al.
Published: (2025)
SEC-QA: A Systematic Evaluation Corpus for Financial QA
by: Lai, Viet Dac, et al.
Published: (2024)
by: Lai, Viet Dac, et al.
Published: (2024)
DeepSearchQA: Bridging the Comprehensiveness Gap for Deep Research Agents
by: Gupta, Nikita, et al.
Published: (2026)
by: Gupta, Nikita, et al.
Published: (2026)
BugBlitz-AI: An Intelligent QA Assistant
by: Yao, Yi, et al.
Published: (2024)
by: Yao, Yi, et al.
Published: (2024)
DocVideoQA: Towards Comprehensive Understanding of Document-Centric Videos through Question Answering
by: Wang, Haochen, et al.
Published: (2025)
by: Wang, Haochen, et al.
Published: (2025)
DisastQA: A Comprehensive Benchmark for Evaluating Question Answering in Disaster Management
by: Chen, Zhitong, et al.
Published: (2026)
by: Chen, Zhitong, et al.
Published: (2026)
LiteraryQA: Towards Effective Evaluation of Long-document Narrative QA
by: Bonomo, Tommaso, et al.
Published: (2025)
by: Bonomo, Tommaso, et al.
Published: (2025)
HumMusQA: A Human-written Music Understanding QA Benchmark Dataset
by: Weck, Benno, et al.
Published: (2026)
by: Weck, Benno, et al.
Published: (2026)
Liver Fibrosis Quantification and Analysis: The LiQA Dataset and Baseline Method
by: Liu, Yuanye, et al.
Published: (2025)
by: Liu, Yuanye, et al.
Published: (2025)
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
by: Zuo, Yuxin, et al.
Published: (2025)
by: Zuo, Yuxin, et al.
Published: (2025)
Compound-QA: A Benchmark for Evaluating LLMs on Compound Questions
by: Hou, Yutao, et al.
Published: (2024)
by: Hou, Yutao, et al.
Published: (2024)
FunQA: Towards Surprising Video Comprehension
by: Xie, Binzhu, et al.
Published: (2023)
by: Xie, Binzhu, et al.
Published: (2023)
ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models
by: Wang, Yueqian, et al.
Published: (2025)
by: Wang, Yueqian, et al.
Published: (2025)
Synthetic Data-Driven Prompt Tuning for Financial QA over Tables and Documents
by: Yu, Yaoning, et al.
Published: (2025)
by: Yu, Yaoning, et al.
Published: (2025)
Selective QA over Conflicting Multi-Source Personal Memory: A Diagnostic Testbed and Method Comparison
by: Yang, Tiancheng, et al.
Published: (2026)
by: Yang, Tiancheng, et al.
Published: (2026)
PolQA: Polish Question Answering Dataset
by: Rybak, Piotr, et al.
Published: (2022)
by: Rybak, Piotr, et al.
Published: (2022)
KET-QA: A Dataset for Knowledge Enhanced Table Question Answering
by: Hu, Mengkang, et al.
Published: (2024)
by: Hu, Mengkang, et al.
Published: (2024)
DeQA-Doc: Adapting DeQA-Score to Document Image Quality Assessment
by: Gao, Junjie, et al.
Published: (2025)
by: Gao, Junjie, et al.
Published: (2025)
ProMQA-Assembly: Multimodal Procedural QA Dataset on Assembly
by: Hasegawa, Kimihiro, et al.
Published: (2025)
by: Hasegawa, Kimihiro, et al.
Published: (2025)
IntentionQA: A Benchmark for Evaluating Purchase Intention Comprehension Abilities of Language Models in E-commerce
by: Ding, Wenxuan, et al.
Published: (2024)
by: Ding, Wenxuan, et al.
Published: (2024)
ECG-Expert-QA: A Benchmark for Evaluating Medical Large Language Models in Heart Disease Diagnosis
by: Wang, Xu, et al.
Published: (2025)
by: Wang, Xu, et al.
Published: (2025)
KGARevion: An AI Agent for Knowledge-Intensive Biomedical QA
by: Su, Xiaorui, et al.
Published: (2024)
by: Su, Xiaorui, et al.
Published: (2024)
GRS-QA -- Graph Reasoning-Structured Question Answering Dataset
by: Pahilajani, Anish, et al.
Published: (2024)
by: Pahilajani, Anish, et al.
Published: (2024)
Advancing 3D Scene Understanding with MV-ScanQA Multi-View Reasoning Evaluation and TripAlign Pre-training Dataset
by: Mo, Wentao, et al.
Published: (2025)
by: Mo, Wentao, et al.
Published: (2025)
VlogQA: Task, Dataset, and Baseline Models for Vietnamese Spoken-Based Machine Reading Comprehension
by: Ngo, Thinh Phuoc, et al.
Published: (2024)
by: Ngo, Thinh Phuoc, et al.
Published: (2024)
SustainableQA: A Comprehensive Question Answering Dataset for Corporate Sustainability and EU Taxonomy Reporting
by: Ali, Mohammed, et al.
Published: (2025)
by: Ali, Mohammed, et al.
Published: (2025)
SimpleDevQA: Benchmarking Large Language Models on Development Knowledge QA
by: Zhang, Jing, et al.
Published: (2025)
by: Zhang, Jing, et al.
Published: (2025)
RepoQA: Evaluating Long Context Code Understanding
by: Liu, Jiawei, et al.
Published: (2024)
by: Liu, Jiawei, et al.
Published: (2024)
FoQA: A Faroese Question-Answering Dataset
by: Simonsen, Annika, et al.
Published: (2025)
by: Simonsen, Annika, et al.
Published: (2025)
QA-LIGN: Aligning LLMs through Constitutionally Decomposed QA
by: Dineen, Jacob, et al.
Published: (2025)
by: Dineen, Jacob, et al.
Published: (2025)
Similar Items
-
RJUA-QA: A Comprehensive QA Dataset for Urology
by: Lyu, Shiwei, et al.
Published: (2023) -
MoleculeQA: A Dataset to Evaluate Factual Accuracy in Molecular Comprehension
by: Lu, Xingyu, et al.
Published: (2024) -
ViQA-COVID: COVID-19 Machine Reading Comprehension Dataset for Vietnamese
by: Nguyen-Phung, Hai-Chung, et al.
Published: (2025) -
JBE-QA: Japanese Bar Exam QA Dataset for Assessing Legal Domain Knowledge
by: Cao, Zhihan, et al.
Published: (2025) -
AgriQA Dataset
by: Eldem, Ayşe
Published: (2026)