Proof of Time: A Benchmark for Evaluating Scientific Idea Judgments

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ye, Bingyang, Chen, Shan, Tu, Jingxuan, Liu, Chen, Xiong, Zidi, Schmidgall, Samuel, Bitterman, Danielle S.
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908759984439296
author Ye, Bingyang
Chen, Shan
Tu, Jingxuan
Liu, Chen
Xiong, Zidi
Schmidgall, Samuel
Bitterman, Danielle S.
author_facet Ye, Bingyang
Chen, Shan
Tu, Jingxuan
Liu, Chen
Xiong, Zidi
Schmidgall, Samuel
Bitterman, Danielle S.
contents Large language models are increasingly being used to assess and forecast research ideas, yet we lack scalable ways to evaluate the quality of models' judgments about these scientific ideas. Towards this goal, we introduce PoT, a semi-verifiable benchmarking framework that links scientific idea judgments to downstream signals that become observable later (e.g., citations and shifts in researchers' agendas). PoT freezes a pre-cutoff snapshot of evidence in an offline sandbox and asks models to forecast post-cutoff outcomes, enabling verifiable evaluation when ground truth arrives, scalable benchmarking without exhaustive expert annotation, and analysis of human-model misalignment against signals such as peer-review awards. In addition, PoT provides a controlled testbed for agent-based research judgments that evaluate scientific ideas, comparing tool-using agents to non-agent baselines under prompt ablations and budget scaling. Across 30,000+ instances spanning four benchmark domains, we find that, compared with non-agent baselines, higher interaction budgets generally improve agent performance, while the benefit of tool use is strongly task-dependent. By combining time-partitioned, future-verifiable targets with an offline sandbox for tool use, PoT supports scalable evaluation of agents on future-facing scientific idea judgment tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2601_07606
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Proof of Time: A Benchmark for Evaluating Scientific Idea Judgments
Ye, Bingyang
Chen, Shan
Tu, Jingxuan
Liu, Chen
Xiong, Zidi
Schmidgall, Samuel
Bitterman, Danielle S.
Computation and Language
Artificial Intelligence
Large language models are increasingly being used to assess and forecast research ideas, yet we lack scalable ways to evaluate the quality of models' judgments about these scientific ideas. Towards this goal, we introduce PoT, a semi-verifiable benchmarking framework that links scientific idea judgments to downstream signals that become observable later (e.g., citations and shifts in researchers' agendas). PoT freezes a pre-cutoff snapshot of evidence in an offline sandbox and asks models to forecast post-cutoff outcomes, enabling verifiable evaluation when ground truth arrives, scalable benchmarking without exhaustive expert annotation, and analysis of human-model misalignment against signals such as peer-review awards. In addition, PoT provides a controlled testbed for agent-based research judgments that evaluate scientific ideas, comparing tool-using agents to non-agent baselines under prompt ablations and budget scaling. Across 30,000+ instances spanning four benchmark domains, we find that, compared with non-agent baselines, higher interaction budgets generally improve agent performance, while the benefit of tool use is strongly task-dependent. By combining time-partitioned, future-verifiable targets with an offline sandbox for tool use, PoT supports scalable evaluation of agents on future-facing scientific idea judgment tasks.
title Proof of Time: A Benchmark for Evaluating Scientific Idea Judgments
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2601.07606