DeepReviewer 2.0: A Traceable Agentic System for Auditable Scientific Peer Review
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911583921242112 |
|---|---|
| author | Weng, Yixuan Zhu, Minjun Xie, Qiujie Ning, Zhiyuan Li, Shichen Lu, Panzhong Lin, Zhen Gu, Enhao Sun, Qiyao Zhang, Yue |
| author_facet | Weng, Yixuan Zhu, Minjun Xie, Qiujie Ning, Zhiyuan Li, Shichen Lu, Panzhong Lin, Zhen Gu, Enhao Sun, Qiyao Zhang, Yue |
| contents | Automated peer review is often framed as generating fluent critique, yet reviewers and area chairs need judgments they can \emph{audit}: where a concern applies, what evidence supports it, and what concrete follow-up is required. DeepReviewer~2.0 is a process-controlled agentic review system built around an output contract: it produces a \textbf{traceable review package} with anchored annotations, localized evidence, and executable follow-up actions, and it exports only after meeting minimum traceability and coverage budgets. Concretely, it first builds a manuscript-only claim--evidence--risk ledger and verification agenda, then performs agenda-driven retrieval and writes anchored critiques under an export gate. On 134 ICLR~2025 submissions under three fixed protocols, an \emph{un-finetuned 196B} model running DeepReviewer~2.0 outperforms Gemini-3.1-Pro-preview, improving strict major-issue coverage (37.26\% vs.\ 23.57\%) and winning 71.63\% of micro-averaged blind comparisons against a human review committee, while ranking first among automatic systems in our pool. We position DeepReviewer~2.0 as an assistive tool rather than a decision proxy, and note remaining gaps such as ethics-sensitive checks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_09590 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | DeepReviewer 2.0: A Traceable Agentic System for Auditable Scientific Peer Review Weng, Yixuan Zhu, Minjun Xie, Qiujie Ning, Zhiyuan Li, Shichen Lu, Panzhong Lin, Zhen Gu, Enhao Sun, Qiyao Zhang, Yue Artificial Intelligence Computation and Language Computers and Society Automated peer review is often framed as generating fluent critique, yet reviewers and area chairs need judgments they can \emph{audit}: where a concern applies, what evidence supports it, and what concrete follow-up is required. DeepReviewer~2.0 is a process-controlled agentic review system built around an output contract: it produces a \textbf{traceable review package} with anchored annotations, localized evidence, and executable follow-up actions, and it exports only after meeting minimum traceability and coverage budgets. Concretely, it first builds a manuscript-only claim--evidence--risk ledger and verification agenda, then performs agenda-driven retrieval and writes anchored critiques under an export gate. On 134 ICLR~2025 submissions under three fixed protocols, an \emph{un-finetuned 196B} model running DeepReviewer~2.0 outperforms Gemini-3.1-Pro-preview, improving strict major-issue coverage (37.26\% vs.\ 23.57\%) and winning 71.63\% of micro-averaged blind comparisons against a human review committee, while ranking first among automatic systems in our pool. We position DeepReviewer~2.0 as an assistive tool rather than a decision proxy, and note remaining gaps such as ethics-sensitive checks. |
| title | DeepReviewer 2.0: A Traceable Agentic System for Auditable Scientific Peer Review |
| topic | Artificial Intelligence Computation and Language Computers and Society |
| url | https://arxiv.org/abs/2604.09590 |