FanOutQA: A Multi-Hop, Multi-Document Question Answering Benchmark for Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhu, Andrew, Hwang, Alyssa, Dugan, Liam, Callison-Burch, Chris |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ReDel: A Toolkit for LLM-Powered Recursive Multi-Agent Systems
von: Zhu, Andrew, et al.
Veröffentlicht: (2024)
von: Zhu, Andrew, et al.
Veröffentlicht: (2024)
Overhearing LLM Agents: A Survey, Taxonomy, and Roadmap
von: Zhu, Andrew, et al.
Veröffentlicht: (2025)
von: Zhu, Andrew, et al.
Veröffentlicht: (2025)
MedHopQA: A Disease-Centered Multi-Hop Reasoning Benchmark and Evaluation Framework for LLM-Based Biomedical Question Answering
von: Islamaj, Rezarta, et al.
Veröffentlicht: (2026)
von: Islamaj, Rezarta, et al.
Veröffentlicht: (2026)
You Have Thirteen Hours in Which to Solve the Labyrinth: Enhancing AI Game Masters with Function Calling
von: Song, Jaewoo, et al.
Veröffentlicht: (2024)
von: Song, Jaewoo, et al.
Veröffentlicht: (2024)
MizanQA: Benchmarking Large Language Models on Moroccan Legal Question Answering
von: Bahaj, Adil, et al.
Veröffentlicht: (2025)
von: Bahaj, Adil, et al.
Veröffentlicht: (2025)
Autorubric: Unifying Rubric-based LLM Evaluation
von: Rao, Delip, et al.
Veröffentlicht: (2026)
von: Rao, Delip, et al.
Veröffentlicht: (2026)
Multi-Hop Reasoning for Question Answering with Hyperbolic Representations
von: Welz, Simon, et al.
Veröffentlicht: (2025)
von: Welz, Simon, et al.
Veröffentlicht: (2025)
Evaluating Monolingual and Multilingual Large Language Models for Greek Question Answering: The DemosQA Benchmark
von: Mastrokostas, Charalampos, et al.
Veröffentlicht: (2026)
von: Mastrokostas, Charalampos, et al.
Veröffentlicht: (2026)
First Steps Towards Overhearing LLM Agents: A Case Study With Dungeons & Dragons Gameplay
von: Zhu, Andrew, et al.
Veröffentlicht: (2025)
von: Zhu, Andrew, et al.
Veröffentlicht: (2025)
Retrieval-enhanced Knowledge Editing in Language Models for Multi-Hop Question Answering
von: Shi, Yucheng, et al.
Veröffentlicht: (2024)
von: Shi, Yucheng, et al.
Veröffentlicht: (2024)
EfficientRAG: Efficient Retriever for Multi-Hop Question Answering
von: Zhuang, Ziyuan, et al.
Veröffentlicht: (2024)
von: Zhuang, Ziyuan, et al.
Veröffentlicht: (2024)
ThinknCheck: Grounded Claim Verification with Compact, Reasoning-Driven, and Interpretable Models
von: Rao, Delip, et al.
Veröffentlicht: (2026)
von: Rao, Delip, et al.
Veröffentlicht: (2026)
Toward Beginner-Friendly LLMs for Language Learning: Controlling Difficulty in Conversation
von: Jin, Meiqing, et al.
Veröffentlicht: (2025)
von: Jin, Meiqing, et al.
Veröffentlicht: (2025)
Credible Plan-Driven RAG Method for Multi-Hop Question Answering
von: Zhang, Ningning, et al.
Veröffentlicht: (2025)
von: Zhang, Ningning, et al.
Veröffentlicht: (2025)
DPRM: A Dual Implicit Process Reward Model in Multi-Hop Question Answering
von: Wang, Xinyi, et al.
Veröffentlicht: (2025)
von: Wang, Xinyi, et al.
Veröffentlicht: (2025)
PRISM: Agentic Retrieval with LLMs for Multi-Hop Question Answering
von: Nahid, Md Mahadi Hasan, et al.
Veröffentlicht: (2025)
von: Nahid, Md Mahadi Hasan, et al.
Veröffentlicht: (2025)
SensorQA: A Question Answering Benchmark for Daily-Life Monitoring
von: Reichman, Benjamin, et al.
Veröffentlicht: (2025)
von: Reichman, Benjamin, et al.
Veröffentlicht: (2025)
MedExQA: Medical Question Answering Benchmark with Multiple Explanations
von: Kim, Yunsoo, et al.
Veröffentlicht: (2024)
von: Kim, Yunsoo, et al.
Veröffentlicht: (2024)
MDER-DR: Multi-Hop Question Answering with Entity-Centric Summaries
von: Campi, Riccardo, et al.
Veröffentlicht: (2026)
von: Campi, Riccardo, et al.
Veröffentlicht: (2026)
A Method for Multi-Hop Question Answering on Persian Knowledge Graph
von: Ghafouri, Arash, et al.
Veröffentlicht: (2025)
von: Ghafouri, Arash, et al.
Veröffentlicht: (2025)
RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors
von: Dugan, Liam, et al.
Veröffentlicht: (2024)
von: Dugan, Liam, et al.
Veröffentlicht: (2024)
Inference Scaled GraphRAG: Improving Multi Hop Question Answering on Knowledge Graphs
von: Thompson, Travis, et al.
Veröffentlicht: (2025)
von: Thompson, Travis, et al.
Veröffentlicht: (2025)
CompactRAG: Reducing LLM Calls and Token Overhead in Multi-Hop Question Answering
von: Yang, Hao, et al.
Veröffentlicht: (2026)
von: Yang, Hao, et al.
Veröffentlicht: (2026)
An Analysis of Decoding Methods for LLM-based Agents for Faithful Multi-Hop Question Answering
von: Murphy, Alexander, et al.
Veröffentlicht: (2025)
von: Murphy, Alexander, et al.
Veröffentlicht: (2025)
MM-PhyQA: Multimodal Physics Question-Answering With Multi-Image CoT Prompting
von: Anand, Avinash, et al.
Veröffentlicht: (2024)
von: Anand, Avinash, et al.
Veröffentlicht: (2024)
MFORT-QA: Multi-hop Few-shot Open Rich Table Question Answering
von: Guan, Che, et al.
Veröffentlicht: (2024)
von: Guan, Che, et al.
Veröffentlicht: (2024)
M3SciQA: A Multi-Modal Multi-Document Scientific QA Benchmark for Evaluating Foundation Models
von: Li, Chuhan, et al.
Veröffentlicht: (2024)
von: Li, Chuhan, et al.
Veröffentlicht: (2024)
Metacognitive Behavioral Tuning of Large Language Models for Multi-Hop Question Answering
von: Kim, Ik-hwan, et al.
Veröffentlicht: (2026)
von: Kim, Ik-hwan, et al.
Veröffentlicht: (2026)
Domain Gating Ensemble Networks for AI-Generated Text Detection
von: Tripathi, Arihant, et al.
Veröffentlicht: (2025)
von: Tripathi, Arihant, et al.
Veröffentlicht: (2025)
Multi-Agents Based on Large Language Models for Knowledge-based Visual Question Answering
von: Hu, Zhongjian, et al.
Veröffentlicht: (2024)
von: Hu, Zhongjian, et al.
Veröffentlicht: (2024)
Bactrainus: Optimizing Large Language Models for Multi-hop Complex Question Answering Tasks
von: Barati, Iman, et al.
Veröffentlicht: (2025)
von: Barati, Iman, et al.
Veröffentlicht: (2025)
Large Language Models Can Self-Improve At Web Agent Tasks
von: Patel, Ajay, et al.
Veröffentlicht: (2024)
von: Patel, Ajay, et al.
Veröffentlicht: (2024)
When Verification Fails: How Compositionally Infeasible Claims Escape Rejection
von: Liu, Muxin, et al.
Veröffentlicht: (2026)
von: Liu, Muxin, et al.
Veröffentlicht: (2026)
STOC-TOT: Stochastic Tree-of-Thought with Constrained Decoding for Complex Reasoning in Multi-Hop Question Answering
von: Bi, Zhenyu, et al.
Veröffentlicht: (2024)
von: Bi, Zhenyu, et al.
Veröffentlicht: (2024)
EpiQAL: Benchmarking Large Language Models in Epidemiological Question Answering and Reasoning
von: Wei, Mingyang, et al.
Veröffentlicht: (2026)
von: Wei, Mingyang, et al.
Veröffentlicht: (2026)
Multi-TW: Benchmarking Multimodal Models on Traditional Chinese Question Answering in Taiwan
von: Yao, Jui-Ming, et al.
Veröffentlicht: (2025)
von: Yao, Jui-Ming, et al.
Veröffentlicht: (2025)
Autofocus Retrieval: An Effective Pipeline for Multi-Hop Question Answering With Semi-Structured Knowledge
von: Boer, Derian, et al.
Veröffentlicht: (2025)
von: Boer, Derian, et al.
Veröffentlicht: (2025)
RELOOP: Recursive Retrieval with Multi-Hop Reasoner and Planners for Heterogeneous QA
von: Yang, Ruiyi, et al.
Veröffentlicht: (2025)
von: Yang, Ruiyi, et al.
Veröffentlicht: (2025)
Multi-hop Question Answering over Knowledge Graphs using Large Language Models
von: Chakraborty, Abir
Veröffentlicht: (2024)
von: Chakraborty, Abir
Veröffentlicht: (2024)
LLM-MedQA: Enhancing Medical Question Answering through Case Studies in Large Language Models
von: Yang, Hang, et al.
Veröffentlicht: (2024)
von: Yang, Hang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
ReDel: A Toolkit for LLM-Powered Recursive Multi-Agent Systems
von: Zhu, Andrew, et al.
Veröffentlicht: (2024) -
Overhearing LLM Agents: A Survey, Taxonomy, and Roadmap
von: Zhu, Andrew, et al.
Veröffentlicht: (2025) -
MedHopQA: A Disease-Centered Multi-Hop Reasoning Benchmark and Evaluation Framework for LLM-Based Biomedical Question Answering
von: Islamaj, Rezarta, et al.
Veröffentlicht: (2026) -
You Have Thirteen Hours in Which to Solve the Labyrinth: Enhancing AI Game Masters with Function Calling
von: Song, Jaewoo, et al.
Veröffentlicht: (2024) -
MizanQA: Benchmarking Large Language Models on Moroccan Legal Question Answering
von: Bahaj, Adil, et al.
Veröffentlicht: (2025)