Navigating Large-Scale Document Collections: MuDABench for Multi-Document Analytical QA
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Zhanli, Cao, Yixuan, Luo, Lvzhou, Luo, Ping |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DeepRead: Document Structure-Aware Reasoning to Enhance Agentic Search
von: Li, Zhanli, et al.
Veröffentlicht: (2026)
von: Li, Zhanli, et al.
Veröffentlicht: (2026)
Attention with Dependency Parsing Augmentation for Fine-Grained Attribution
von: Ding, Qiang, et al.
Veröffentlicht: (2024)
von: Ding, Qiang, et al.
Veröffentlicht: (2024)
The Gray Zone of Faithfulness: Taming Ambiguity in Unfaithfulness Detection
von: Ding, Qiang, et al.
Veröffentlicht: (2025)
von: Ding, Qiang, et al.
Veröffentlicht: (2025)
ChiMDQA: Towards Comprehensive Chinese Document QA with Fine-grained Evaluation
von: Gao, Jing, et al.
Veröffentlicht: (2025)
von: Gao, Jing, et al.
Veröffentlicht: (2025)
FinDABench: Benchmarking Financial Data Analysis Ability of Large Language Models
von: Liu, Shu, et al.
Veröffentlicht: (2024)
von: Liu, Shu, et al.
Veröffentlicht: (2024)
AttnComp: Attention-Guided Adaptive Context Compression for Retrieval-Augmented Generation
von: Luo, Lvzhou, et al.
Veröffentlicht: (2025)
von: Luo, Lvzhou, et al.
Veröffentlicht: (2025)
Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QA
von: Wang, Minzheng, et al.
Veröffentlicht: (2024)
von: Wang, Minzheng, et al.
Veröffentlicht: (2024)
Reasoning Pattern Matters: Learning to Reason without Human Rationales
von: Pang, Chaoxu, et al.
Veröffentlicht: (2025)
von: Pang, Chaoxu, et al.
Veröffentlicht: (2025)
InfiAgent-DABench: Evaluating Agents on Data Analysis Tasks
von: Hu, Xueyu, et al.
Veröffentlicht: (2024)
von: Hu, Xueyu, et al.
Veröffentlicht: (2024)
M3SciQA: A Multi-Modal Multi-Document Scientific QA Benchmark for Evaluating Foundation Models
von: Li, Chuhan, et al.
Veröffentlicht: (2024)
von: Li, Chuhan, et al.
Veröffentlicht: (2024)
MuLD: The Multitask Long Document Benchmark
von: Hudson, G Thomas, et al.
Veröffentlicht: (2022)
von: Hudson, G Thomas, et al.
Veröffentlicht: (2022)
ConDABench: Interactive Evaluation of Language Models for Data Analysis
von: Dutta, Avik, et al.
Veröffentlicht: (2025)
von: Dutta, Avik, et al.
Veröffentlicht: (2025)
FanOutQA: A Multi-Hop, Multi-Document Question Answering Benchmark for Large Language Models
von: Zhu, Andrew, et al.
Veröffentlicht: (2024)
von: Zhu, Andrew, et al.
Veröffentlicht: (2024)
HiQA: A Hierarchical Contextual Augmentation RAG for Multi-Documents QA
von: Chen, Xinyue, et al.
Veröffentlicht: (2024)
von: Chen, Xinyue, et al.
Veröffentlicht: (2024)
Long-Document QA with Chain-of-Structured-Thought and Fine-Tuned SLMs
von: Liang, Zhuowen, et al.
Veröffentlicht: (2026)
von: Liang, Zhuowen, et al.
Veröffentlicht: (2026)
DesignQA: A Multimodal Benchmark for Evaluating Large Language Models' Understanding of Engineering Documentation
von: Doris, Anna C., et al.
Veröffentlicht: (2024)
von: Doris, Anna C., et al.
Veröffentlicht: (2024)
Strategic Navigation or Stochastic Search? How Agents and Humans Reason Over Document Collections
von: Borchmann, Łukasz, et al.
Veröffentlicht: (2026)
von: Borchmann, Łukasz, et al.
Veröffentlicht: (2026)
AirQA: A Comprehensive QA Dataset for AI Research with Instance-Level Evaluation
von: Huang, Tiancheng, et al.
Veröffentlicht: (2025)
von: Huang, Tiancheng, et al.
Veröffentlicht: (2025)
EchoQA: A Large Collection of Instruction Tuning Data for Echocardiogram Reports
von: Moukheiber, Lama, et al.
Veröffentlicht: (2025)
von: Moukheiber, Lama, et al.
Veröffentlicht: (2025)
D-SCoRE: Document-Centric Segmentation and CoT Reasoning with Structured Export for QA-CoT Data Generation
von: Zhou, Weibo, et al.
Veröffentlicht: (2025)
von: Zhou, Weibo, et al.
Veröffentlicht: (2025)
LLM Augmentations to support Analytical Reasoning over Multiple Documents
von: Yousuf, Raquib Bin, et al.
Veröffentlicht: (2024)
von: Yousuf, Raquib Bin, et al.
Veröffentlicht: (2024)
KunlunBaize: LLM with Multi-Scale Convolution and Multi-Token Prediction Under TransformerX Framework
von: Li, Cheng, et al.
Veröffentlicht: (2025)
von: Li, Cheng, et al.
Veröffentlicht: (2025)
The First Place Solution of WSDM Cup 2024: Leveraging Large Language Models for Conversational Multi-Doc QA
von: Li, Yiming, et al.
Veröffentlicht: (2024)
von: Li, Yiming, et al.
Veröffentlicht: (2024)
Document-Level Event Extraction with Definition-Driven ICL
von: Liu, Zhuoyuan, et al.
Veröffentlicht: (2024)
von: Liu, Zhuoyuan, et al.
Veröffentlicht: (2024)
OCR or Not? Rethinking Document Information Extraction in the MLLMs Era with Real-World Large-Scale Datasets
von: Shen, Jiyuan, et al.
Veröffentlicht: (2026)
von: Shen, Jiyuan, et al.
Veröffentlicht: (2026)
CollectiveSFT: Scaling Large Language Models for Chinese Medical Benchmark with Collective Instructions in Healthcare
von: Zhu, Jingwei, et al.
Veröffentlicht: (2024)
von: Zhu, Jingwei, et al.
Veröffentlicht: (2024)
DABench-LLM: Standardized and In-Depth Benchmarking of Post-Moore Dataflow AI Accelerators for LLMs
von: Hu, Ziyu, et al.
Veröffentlicht: (2025)
von: Hu, Ziyu, et al.
Veröffentlicht: (2025)
MuSC: Improving Complex Instruction Following with Multi-granularity Self-Contrastive Training
von: Huang, Hui, et al.
Veröffentlicht: (2025)
von: Huang, Hui, et al.
Veröffentlicht: (2025)
MuMath-Code: Combining Tool-Use Large Language Models with Multi-perspective Data Augmentation for Mathematical Reasoning
von: Yin, Shuo, et al.
Veröffentlicht: (2024)
von: Yin, Shuo, et al.
Veröffentlicht: (2024)
Large Language Models for Single-Step and Multi-Step Flight Trajectory Prediction
von: Luo, Kaiwei, et al.
Veröffentlicht: (2025)
von: Luo, Kaiwei, et al.
Veröffentlicht: (2025)
Enhancing Legal Document Retrieval: A Multi-Phase Approach with Large Language Models
von: Nguyen, Hai-Long, et al.
Veröffentlicht: (2024)
von: Nguyen, Hai-Long, et al.
Veröffentlicht: (2024)
MuTSE: A Human-in-the-Loop Multi-use Text Simplification Evaluator
von: Roscan, Rares-Alexandru, et al.
Veröffentlicht: (2026)
von: Roscan, Rares-Alexandru, et al.
Veröffentlicht: (2026)
Document Intelligence in the Era of Large Language Models: A Survey
von: Wang, Weishi, et al.
Veröffentlicht: (2025)
von: Wang, Weishi, et al.
Veröffentlicht: (2025)
Memorizing Documents with Guidance in Large Language Models
von: Park, Bumjin, et al.
Veröffentlicht: (2024)
von: Park, Bumjin, et al.
Veröffentlicht: (2024)
Docs2KG: Unified Knowledge Graph Construction from Heterogeneous Documents Assisted by Large Language Models
von: Sun, Qiang, et al.
Veröffentlicht: (2024)
von: Sun, Qiang, et al.
Veröffentlicht: (2024)
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages
von: Han, Wenhan, et al.
Veröffentlicht: (2025)
von: Han, Wenhan, et al.
Veröffentlicht: (2025)
Leveraging Long-Context Large Language Models for Multi-Document Understanding and Summarization in Enterprise Applications
von: Godbole, Aditi, et al.
Veröffentlicht: (2024)
von: Godbole, Aditi, et al.
Veröffentlicht: (2024)
Self-Prompting Large Language Models for Zero-Shot Open-Domain QA
von: Li, Junlong, et al.
Veröffentlicht: (2022)
von: Li, Junlong, et al.
Veröffentlicht: (2022)
Multi-Document Grounded Multi-Turn Synthetic Dialog Generation
von: Lee, Young-Suk, et al.
Veröffentlicht: (2024)
von: Lee, Young-Suk, et al.
Veröffentlicht: (2024)
MedCritical: Enhancing Medical Reasoning in Small Language Models via Self-Collaborative Correction
von: Su, Xinchun, et al.
Veröffentlicht: (2025)
von: Su, Xinchun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
DeepRead: Document Structure-Aware Reasoning to Enhance Agentic Search
von: Li, Zhanli, et al.
Veröffentlicht: (2026) -
Attention with Dependency Parsing Augmentation for Fine-Grained Attribution
von: Ding, Qiang, et al.
Veröffentlicht: (2024) -
The Gray Zone of Faithfulness: Taming Ambiguity in Unfaithfulness Detection
von: Ding, Qiang, et al.
Veröffentlicht: (2025) -
ChiMDQA: Towards Comprehensive Chinese Document QA with Fine-grained Evaluation
von: Gao, Jing, et al.
Veröffentlicht: (2025) -
FinDABench: Benchmarking Financial Data Analysis Ability of Large Language Models
von: Liu, Shu, et al.
Veröffentlicht: (2024)