MoNaCo: More Natural and Complex Questions for Reasoning Across Dozens of Documents

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wolfson, Tomer, Trivedi, Harsh, Geva, Mor, Goldberg, Yoav, Roth, Dan, Khot, Tushar, Sabharwal, Ashish, Tsarfaty, Reut
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915476675756032
author Wolfson, Tomer
Trivedi, Harsh
Geva, Mor
Goldberg, Yoav
Roth, Dan
Khot, Tushar
Sabharwal, Ashish
Tsarfaty, Reut
author_facet Wolfson, Tomer
Trivedi, Harsh
Geva, Mor
Goldberg, Yoav
Roth, Dan
Khot, Tushar
Sabharwal, Ashish
Tsarfaty, Reut
contents Automated agents, powered by Large language models (LLMs), are emerging as the go-to tool for querying information. However, evaluation benchmarks for LLM agents rarely feature natural questions that are both information-seeking and genuinely time-consuming for humans. To address this gap we introduce MoNaCo, a benchmark of 1,315 natural and time-consuming questions that require dozens, and at times hundreds, of intermediate steps to solve -- far more than any existing QA benchmark. To build MoNaCo, we developed a decomposed annotation pipeline to elicit and manually answer real-world time-consuming questions at scale. Frontier LLMs evaluated on MoNaCo achieve at most 61.2% F1, hampered by low recall and hallucinations. Our results underscore the limitations of LLM-powered agents in handling the complexity and sheer breadth of real-world information-seeking tasks -- with MoNaCo providing an effective resource for tracking such progress. The MoNaCo benchmark, codebase, prompts and models predictions are all publicly available at: https://tomerwolgithub.github.io/monaco
format Preprint
id arxiv_https___arxiv_org_abs_2508_11133
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MoNaCo: More Natural and Complex Questions for Reasoning Across Dozens of Documents
Wolfson, Tomer
Trivedi, Harsh
Geva, Mor
Goldberg, Yoav
Roth, Dan
Khot, Tushar
Sabharwal, Ashish
Tsarfaty, Reut
Computation and Language
Artificial Intelligence
Databases
Automated agents, powered by Large language models (LLMs), are emerging as the go-to tool for querying information. However, evaluation benchmarks for LLM agents rarely feature natural questions that are both information-seeking and genuinely time-consuming for humans. To address this gap we introduce MoNaCo, a benchmark of 1,315 natural and time-consuming questions that require dozens, and at times hundreds, of intermediate steps to solve -- far more than any existing QA benchmark. To build MoNaCo, we developed a decomposed annotation pipeline to elicit and manually answer real-world time-consuming questions at scale. Frontier LLMs evaluated on MoNaCo achieve at most 61.2% F1, hampered by low recall and hallucinations. Our results underscore the limitations of LLM-powered agents in handling the complexity and sheer breadth of real-world information-seeking tasks -- with MoNaCo providing an effective resource for tracking such progress. The MoNaCo benchmark, codebase, prompts and models predictions are all publicly available at: https://tomerwolgithub.github.io/monaco
title MoNaCo: More Natural and Complex Questions for Reasoning Across Dozens of Documents
topic Computation and Language
Artificial Intelligence
Databases
url https://arxiv.org/abs/2508.11133