ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Meng, Rui, Mishra, Bhavana Dalvi, Chen, Jiefeng, Li, Chun-Liang, Goyal, Palash, Parmar, Mihir, Song, Yiwen, Song, Yale, Sinha, Rajarishi, Ranganathan, Parthasarathy, Gokturk, Burak, Yoon, Jinsung, Pfister, Tomas
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913162463281152
author Meng, Rui
Mishra, Bhavana Dalvi
Chen, Jiefeng
Li, Chun-Liang
Goyal, Palash
Parmar, Mihir
Song, Yiwen
Song, Yale
Sinha, Rajarishi
Ranganathan, Parthasarathy
Gokturk, Burak
Yoon, Jinsung
Pfister, Tomas
author_facet Meng, Rui
Mishra, Bhavana Dalvi
Chen, Jiefeng
Li, Chun-Liang
Goyal, Palash
Parmar, Mihir
Song, Yiwen
Song, Yale
Sinha, Rajarishi
Ranganathan, Parthasarathy
Gokturk, Burak
Yoon, Jinsung
Pfister, Tomas
contents Autonomous research agents produce competitive solutions and professional-looking manuscripts, yet their outputs contain verifiability failures undetectable by surface-level evaluation: fabricated citations, unreproducible scores, and method descriptions that diverge from the implementation. We address this through three contributions. First, Chain-of-Evidence (CoE), a verifiability framework requiring every claim to be traceable to its evidence source. Second, ScientistOne, an end-to-end autonomous research system that maintains evidence chains by construction throughout literature review, solution discovery, and paper writing. Third, CoE Audit, a post-hoc audit whose four integrity checks -- score verification, specification violation, reference verification, and method-code alignment -- apply uniformly to all systems. Across 75 papers spanning five systems and five frontier research tasks, every baseline exhibits at least one systematic failure mode: hallucinated reference rates reach 21%, score verification passes in as few as 42% of papers, and method-code alignment ranges from 20% to 80%. ScientistOne achieves zero hallucinated references (0/337), perfect score verification (12/12), and the highest method-code alignment (14/15), while matching or exceeding human expert performance on all five tasks. ScientistOne further generalizes to six additional tasks spanning medical imaging, fine-grained recognition, 3D perception, and language modeling, achieving state-of-the-art on Parameter Golf and gold medals on MLE-Bench tasks where baselines fail entirely.
format Preprint
id arxiv_https___arxiv_org_abs_2605_26340
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence
Meng, Rui
Mishra, Bhavana Dalvi
Chen, Jiefeng
Li, Chun-Liang
Goyal, Palash
Parmar, Mihir
Song, Yiwen
Song, Yale
Sinha, Rajarishi
Ranganathan, Parthasarathy
Gokturk, Burak
Yoon, Jinsung
Pfister, Tomas
Artificial Intelligence
Computation and Language
Multiagent Systems
Autonomous research agents produce competitive solutions and professional-looking manuscripts, yet their outputs contain verifiability failures undetectable by surface-level evaluation: fabricated citations, unreproducible scores, and method descriptions that diverge from the implementation. We address this through three contributions. First, Chain-of-Evidence (CoE), a verifiability framework requiring every claim to be traceable to its evidence source. Second, ScientistOne, an end-to-end autonomous research system that maintains evidence chains by construction throughout literature review, solution discovery, and paper writing. Third, CoE Audit, a post-hoc audit whose four integrity checks -- score verification, specification violation, reference verification, and method-code alignment -- apply uniformly to all systems. Across 75 papers spanning five systems and five frontier research tasks, every baseline exhibits at least one systematic failure mode: hallucinated reference rates reach 21%, score verification passes in as few as 42% of papers, and method-code alignment ranges from 20% to 80%. ScientistOne achieves zero hallucinated references (0/337), perfect score verification (12/12), and the highest method-code alignment (14/15), while matching or exceeding human expert performance on all five tasks. ScientistOne further generalizes to six additional tasks spanning medical imaging, fine-grained recognition, 3D perception, and language modeling, achieving state-of-the-art on Parameter Golf and gold medals on MLE-Bench tasks where baselines fail entirely.
title ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence
topic Artificial Intelligence
Computation and Language
Multiagent Systems
url https://arxiv.org/abs/2605.26340