MoNaCo: More Natural and Complex Questions for Reasoning Across Dozens of Documents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866915476675756032 |
|---|---|
| author | Wolfson, Tomer Trivedi, Harsh Geva, Mor Goldberg, Yoav Roth, Dan Khot, Tushar Sabharwal, Ashish Tsarfaty, Reut |
| author_facet | Wolfson, Tomer Trivedi, Harsh Geva, Mor Goldberg, Yoav Roth, Dan Khot, Tushar Sabharwal, Ashish Tsarfaty, Reut |
| contents | Automated agents, powered by Large language models (LLMs), are emerging as the go-to tool for querying information. However, evaluation benchmarks for LLM agents rarely feature natural questions that are both information-seeking and genuinely time-consuming for humans. To address this gap we introduce MoNaCo, a benchmark of 1,315 natural and time-consuming questions that require dozens, and at times hundreds, of intermediate steps to solve -- far more than any existing QA benchmark. To build MoNaCo, we developed a decomposed annotation pipeline to elicit and manually answer real-world time-consuming questions at scale. Frontier LLMs evaluated on MoNaCo achieve at most 61.2% F1, hampered by low recall and hallucinations. Our results underscore the limitations of LLM-powered agents in handling the complexity and sheer breadth of real-world information-seeking tasks -- with MoNaCo providing an effective resource for tracking such progress. The MoNaCo benchmark, codebase, prompts and models predictions are all publicly available at: https://tomerwolgithub.github.io/monaco |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_11133 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | MoNaCo: More Natural and Complex Questions for Reasoning Across Dozens of Documents Wolfson, Tomer Trivedi, Harsh Geva, Mor Goldberg, Yoav Roth, Dan Khot, Tushar Sabharwal, Ashish Tsarfaty, Reut Computation and Language Artificial Intelligence Databases Automated agents, powered by Large language models (LLMs), are emerging as the go-to tool for querying information. However, evaluation benchmarks for LLM agents rarely feature natural questions that are both information-seeking and genuinely time-consuming for humans. To address this gap we introduce MoNaCo, a benchmark of 1,315 natural and time-consuming questions that require dozens, and at times hundreds, of intermediate steps to solve -- far more than any existing QA benchmark. To build MoNaCo, we developed a decomposed annotation pipeline to elicit and manually answer real-world time-consuming questions at scale. Frontier LLMs evaluated on MoNaCo achieve at most 61.2% F1, hampered by low recall and hallucinations. Our results underscore the limitations of LLM-powered agents in handling the complexity and sheer breadth of real-world information-seeking tasks -- with MoNaCo providing an effective resource for tracking such progress. The MoNaCo benchmark, codebase, prompts and models predictions are all publicly available at: https://tomerwolgithub.github.io/monaco |
| title | MoNaCo: More Natural and Complex Questions for Reasoning Across Dozens of Documents |
| topic | Computation and Language Artificial Intelligence Databases |
| url | https://arxiv.org/abs/2508.11133 |