pdfQA: Diverse, Challenging, and Realistic Question Answering over PDFs
Fuente:
arXiv
Saved in:
| Main Authors: | Schimanski, Tobias, Kolli, Imene, Fan, Yu, Vaghefi, Ario Saeid, Ni, Jingwei, Ash, Elliott, Leippold, Markus |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Faithful and Robust LLM Specialists for Evidence-Based Question-Answering
by: Schimanski, Tobias, et al.
Published: (2024)
by: Schimanski, Tobias, et al.
Published: (2024)
Automated Evidence Extraction and Scoring for Corporate Climate Policy Engagement: A Multilingual RAG Approach
by: Kolli, Imene, et al.
Published: (2025)
by: Kolli, Imene, et al.
Published: (2025)
DIRAS: Efficient LLM Annotation of Document Relevance in Retrieval Augmented Generation
by: Ni, Jingwei, et al.
Published: (2024)
by: Ni, Jingwei, et al.
Published: (2024)
ClimRetrieve: A Benchmarking Dataset for Information Retrieval from Corporate Climate Disclosures
by: Schimanski, Tobias, et al.
Published: (2024)
by: Schimanski, Tobias, et al.
Published: (2024)
AFaCTA: Assisting the Annotation of Factual Claim Detection with Reliable LLM Annotators
by: Ni, Jingwei, et al.
Published: (2024)
by: Ni, Jingwei, et al.
Published: (2024)
AI for Climate Finance: Agentic Retrieval and Multi-Step Reasoning for Early Warning System Investments
by: Vaghefi, Saeid Ario, et al.
Published: (2025)
by: Vaghefi, Saeid Ario, et al.
Published: (2025)
Exploring Nature: Datasets and Models for Analyzing Nature-Related Disclosures
by: Schimanski, Tobias, et al.
Published: (2023)
by: Schimanski, Tobias, et al.
Published: (2023)
Automated Fact-Checking of Climate Change Claims with Large Language Models
by: Leippold, Markus, et al.
Published: (2024)
by: Leippold, Markus, et al.
Published: (2024)
Can Reasoning Help Large Language Models Capture Human Annotator Disagreement?
by: Ni, Jingwei, et al.
Published: (2025)
by: Ni, Jingwei, et al.
Published: (2025)
Balancing Truthfulness and Informativeness with Uncertainty-Aware Instruction Fine-Tuning
by: Wu, Tianyi, et al.
Published: (2025)
by: Wu, Tianyi, et al.
Published: (2025)
UsefulBench: Towards Decision-Useful Information as a Target for Information Retrieval
by: Schimanski, Tobias, et al.
Published: (2026)
by: Schimanski, Tobias, et al.
Published: (2026)
Harmonized World Soil Database in SWAT Format
by: Abbaspour, Karim, et al.
Published: (2019)
by: Abbaspour, Karim, et al.
Published: (2019)
Global Land Cover for SWAT "Global Landuse GlobCover "
by: Abbaspour, Karim, et al.
Published: (2019)
by: Abbaspour, Karim, et al.
Published: (2019)
CRU and GCM data for SWAT model
by: Abbaspour, Karim, et al.
Published: (2019)
by: Abbaspour, Karim, et al.
Published: (2019)
Global Land Cover for SWAT "Global Landuse USGS"
by: Abbaspour, Karim, et al.
Published: (2019)
by: Abbaspour, Karim, et al.
Published: (2019)
Global FAO/UNESCO Soil Map of the World reformatted with SWAT format
by: Abbaspour, Karim, et al.
Published: (2019)
by: Abbaspour, Karim, et al.
Published: (2019)
ReProbe: Efficient Test-Time Scaling of Multi-Step Reasoning by Probing Internal States of Large Language Models
by: Ni, Jingwei, et al.
Published: (2025)
by: Ni, Jingwei, et al.
Published: (2025)
ASTRA-QA: A Benchmark for Abstract Question Answering over Documents
by: Wang, Shu, et al.
Published: (2026)
by: Wang, Shu, et al.
Published: (2026)
Co-DETECT: Collaborative Discovery of Edge Cases in Text Classification
by: Xiong, Chenfei, et al.
Published: (2025)
by: Xiong, Chenfei, et al.
Published: (2025)
NeoQA: Evidence-based Question Answering with Generated News Events
by: Glockner, Max, et al.
Published: (2025)
by: Glockner, Max, et al.
Published: (2025)
Good Question! The Effect of Positive Feedback on Contributions to Online Public Goods
by: Wachs, Johannes, et al.
Published: (2026)
by: Wachs, Johannes, et al.
Published: (2026)
FHIRPath-QA: Executable Question Answering over FHIR Electronic Health Records
by: Frew, Michael, et al.
Published: (2026)
by: Frew, Michael, et al.
Published: (2026)
PolQA: Polish Question Answering Dataset
by: Rybak, Piotr, et al.
Published: (2022)
by: Rybak, Piotr, et al.
Published: (2022)
VoQA: Visual-only Question Answering
by: An, Jianing, et al.
Published: (2025)
by: An, Jianing, et al.
Published: (2025)
Word-Centered Semantic Graphs for Interpretable Diachronic Sense Tracking
by: Kolli, Imene, et al.
Published: (2026)
by: Kolli, Imene, et al.
Published: (2026)
CAN-QA: A Question-Answering Benchmark for Reasoning over In-Vehicle CAN Traffic
by: Chen, Jing, et al.
Published: (2026)
by: Chen, Jing, et al.
Published: (2026)
WikiMixQA: A Multimodal Benchmark for Question Answering over Tables and Charts
by: Foroutan, Negar, et al.
Published: (2025)
by: Foroutan, Negar, et al.
Published: (2025)
GRS-QA -- Graph Reasoning-Structured Question Answering Dataset
by: Pahilajani, Anish, et al.
Published: (2024)
by: Pahilajani, Anish, et al.
Published: (2024)
LingoQA: Visual Question Answering for Autonomous Driving
by: Marcu, Ana-Maria, et al.
Published: (2023)
by: Marcu, Ana-Maria, et al.
Published: (2023)
Answering Questions in Stages: Prompt Chaining for Contract QA
by: Roegiest, Adam, et al.
Published: (2024)
by: Roegiest, Adam, et al.
Published: (2024)
DebateQA: Evaluating Question Answering on Debatable Knowledge
by: Xu, Rongwu, et al.
Published: (2024)
by: Xu, Rongwu, et al.
Published: (2024)
FoQA: A Faroese Question-Answering Dataset
by: Simonsen, Annika, et al.
Published: (2025)
by: Simonsen, Annika, et al.
Published: (2025)
CT2C-QA: Multimodal Question Answering over Chinese Text, Table and Chart
by: Zhao, Bowen, et al.
Published: (2024)
by: Zhao, Bowen, et al.
Published: (2024)
MA-EgoQA: Question Answering over Egocentric Videos from Multiple Embodied Agents
by: Kim, Kangsan, et al.
Published: (2026)
by: Kim, Kangsan, et al.
Published: (2026)
Beyond Factual QA: Mentorship-Oriented Question Answering over Long-Form Multilingual Content
by: Bhalerao, Parth, et al.
Published: (2026)
by: Bhalerao, Parth, et al.
Published: (2026)
SensorQA: A Question Answering Benchmark for Daily-Life Monitoring
by: Reichman, Benjamin, et al.
Published: (2025)
by: Reichman, Benjamin, et al.
Published: (2025)
ClimaQA: An Automated Evaluation Framework for Climate Question Answering Models
by: Manivannan, Veeramakali Vignesh, et al.
Published: (2024)
by: Manivannan, Veeramakali Vignesh, et al.
Published: (2024)
ClimaQA_SLO - Slovenian Climate Question-Answering Benchmark
by: Ferk Ovčjak, Monika, et al.
Published: (2025)
by: Ferk Ovčjak, Monika, et al.
Published: (2025)
MMToM-QA: Multimodal Theory of Mind Question Answering
by: Jin, Chuanyang, et al.
Published: (2024)
by: Jin, Chuanyang, et al.
Published: (2024)
SyllabusQA: A Course Logistics Question Answering Dataset
by: Fernandez, Nigel, et al.
Published: (2024)
by: Fernandez, Nigel, et al.
Published: (2024)
Similar Items
-
Towards Faithful and Robust LLM Specialists for Evidence-Based Question-Answering
by: Schimanski, Tobias, et al.
Published: (2024) -
Automated Evidence Extraction and Scoring for Corporate Climate Policy Engagement: A Multilingual RAG Approach
by: Kolli, Imene, et al.
Published: (2025) -
DIRAS: Efficient LLM Annotation of Document Relevance in Retrieval Augmented Generation
by: Ni, Jingwei, et al.
Published: (2024) -
ClimRetrieve: A Benchmarking Dataset for Information Retrieval from Corporate Climate Disclosures
by: Schimanski, Tobias, et al.
Published: (2024) -
AFaCTA: Assisting the Annotation of Factual Claim Detection with Reliable LLM Annotators
by: Ni, Jingwei, et al.
Published: (2024)