Hunt Instead of Wait: Evaluating Deep Data Research on Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Wei, Yu, Peijie, Orini, Michele, Du, Yali, He, Yulan
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916017257578496
author Liu, Wei
Yu, Peijie
Orini, Michele
Du, Yali
He, Yulan
author_facet Liu, Wei
Yu, Peijie
Orini, Michele
Du, Yali
He, Yulan
contents The agency expected of Agentic Large Language Models goes beyond answering correctly, requiring autonomy to set goals and decide what to explore. We term this investigatory intelligence, distinguishing it from executional intelligence, which merely completes assigned tasks. Data Science provides a natural testbed, as real-world analysis starts from raw data rather than explicit queries, yet few benchmarks focus on it. To address this, we introduce Deep Data Research (DDR), an open-ended task where LLMs autonomously extract key insights from databases, and DDR-Bench, a large-scale, checklist-based benchmark that enables verifiable evaluation. Results show that while frontier models display emerging agency, long-horizon exploration remains challenging. Our analysis highlights that effective investigatory intelligence depends not only on agent scaffolding or merely scaling, but also on intrinsic strategies of agentic models.
format Preprint
id arxiv_https___arxiv_org_abs_2602_02039
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Hunt Instead of Wait: Evaluating Deep Data Research on Large Language Models
Liu, Wei
Yu, Peijie
Orini, Michele
Du, Yali
He, Yulan
Artificial Intelligence
Computation and Language
Databases
Machine Learning
The agency expected of Agentic Large Language Models goes beyond answering correctly, requiring autonomy to set goals and decide what to explore. We term this investigatory intelligence, distinguishing it from executional intelligence, which merely completes assigned tasks. Data Science provides a natural testbed, as real-world analysis starts from raw data rather than explicit queries, yet few benchmarks focus on it. To address this, we introduce Deep Data Research (DDR), an open-ended task where LLMs autonomously extract key insights from databases, and DDR-Bench, a large-scale, checklist-based benchmark that enables verifiable evaluation. Results show that while frontier models display emerging agency, long-horizon exploration remains challenging. Our analysis highlights that effective investigatory intelligence depends not only on agent scaffolding or merely scaling, but also on intrinsic strategies of agentic models.
title Hunt Instead of Wait: Evaluating Deep Data Research on Large Language Models
topic Artificial Intelligence
Computation and Language
Databases
Machine Learning
url https://arxiv.org/abs/2602.02039