PreScience: A Benchmark for Forecasting Scientific Contributions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ajith, Anirudh, Singh, Amanpreet, DeYoung, Jay, Kunievsky, Nadav, Kozlowski, Austin C., Tafjord, Oyvind, Evans, James, Weld, Daniel S., Hope, Tom, Downey, Doug |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The (Short-Term) Effects of Large Language Models on Unemployment and Earnings
von: Chen, Danqing, et al.
Veröffentlicht: (2025)
von: Chen, Danqing, et al.
Veröffentlicht: (2025)
SciMON: Scientific Inspiration Machines Optimized for Novelty
von: Wang, Qingyun, et al.
Veröffentlicht: (2023)
von: Wang, Qingyun, et al.
Veröffentlicht: (2023)
Generating Literature-Driven Scientific Theories at Scale
von: Jansen, Peter, et al.
Veröffentlicht: (2026)
von: Jansen, Peter, et al.
Veröffentlicht: (2026)
Polarization by Design: How Elites Could Shape Mass Preferences as AI Reduces Persuasion Costs
von: Kunievsky, Nadav
Veröffentlicht: (2025)
von: Kunievsky, Nadav
Veröffentlicht: (2025)
Linear Regression in a Nonlinear World
von: Kunievsky, Nadav
Veröffentlicht: (2025)
von: Kunievsky, Nadav
Veröffentlicht: (2025)
The Effect of Age at Arrival on the Alignment Between Immigrant and Native-Born Gender Norms: A Distributional Approach
von: Kunievsky, Nadav
Veröffentlicht: (2026)
von: Kunievsky, Nadav
Veröffentlicht: (2026)
MARG: Multi-Agent Review Generation for Scientific Papers
von: D'Arcy, Mike, et al.
Veröffentlicht: (2024)
von: D'Arcy, Mike, et al.
Veröffentlicht: (2024)
Measuring Intent Comprehension in LLMs
von: Kunievsky, Nadav, et al.
Veröffentlicht: (2025)
von: Kunievsky, Nadav, et al.
Veröffentlicht: (2025)
CodeScientist: End-to-End Semi-Automated Scientific Discovery with Code-based Experimentation
von: Jansen, Peter, et al.
Veröffentlicht: (2025)
von: Jansen, Peter, et al.
Veröffentlicht: (2025)
Content vs. Form: What Drives the Writing Score Gap Across Socioeconomic Backgrounds? A Generated Panel Approach
von: Kunievsky, Nadav, et al.
Veröffentlicht: (2026)
von: Kunievsky, Nadav, et al.
Veröffentlicht: (2026)
Missing vs. Unused Knowledge Hypothesis for Language Model Bottlenecks in Patent Understanding
von: Wu, Siyang, et al.
Veröffentlicht: (2025)
von: Wu, Siyang, et al.
Veröffentlicht: (2025)
ARIES: A Corpus of Scientific Paper Edits Made in Response to Peer Reviews
von: D'Arcy, Mike, et al.
Veröffentlicht: (2023)
von: D'Arcy, Mike, et al.
Veröffentlicht: (2023)
CARE: Extracting Experimental Findings From Clinical Literature
von: Naik, Aakanksha, et al.
Veröffentlicht: (2023)
von: Naik, Aakanksha, et al.
Veröffentlicht: (2023)
Literature-Grounded Novelty Assessment of Scientific Ideas
von: Shahid, Simra, et al.
Veröffentlicht: (2025)
von: Shahid, Simra, et al.
Veröffentlicht: (2025)
Biased AI improves human decision-making but reduces trust
von: Lai, Shiyang, et al.
Veröffentlicht: (2025)
von: Lai, Shiyang, et al.
Veröffentlicht: (2025)
Digital Socrates: Evaluating LLMs through Explanation Critiques
von: Gu, Yuling, et al.
Veröffentlicht: (2023)
von: Gu, Yuling, et al.
Veröffentlicht: (2023)
TOPICAL: TOPIC Pages AutomagicaLly
von: Giorgi, John, et al.
Veröffentlicht: (2024)
von: Giorgi, John, et al.
Veröffentlicht: (2024)
Omakase: proactive assistance with actionable suggestions for evolving scientific research projects
von: Siangliulue, Pao, et al.
Veröffentlicht: (2026)
von: Siangliulue, Pao, et al.
Veröffentlicht: (2026)
Human-LLM Compound System for Scientific Ideation through Facet Recombination and Novelty Evaluation
von: Radensky, Marissa, et al.
Veröffentlicht: (2024)
von: Radensky, Marissa, et al.
Veröffentlicht: (2024)
In Silico Sociology: Forecasting COVID-19 Polarization with Large Language Models
von: Kozlowski, Austin C., et al.
Veröffentlicht: (2024)
von: Kozlowski, Austin C., et al.
Veröffentlicht: (2024)
BaRDa: A Belief and Reasoning Dataset that Separates Factual Accuracy and Reasoning Ability
von: Clark, Peter, et al.
Veröffentlicht: (2023)
von: Clark, Peter, et al.
Veröffentlicht: (2023)
Deep Research, Shallow Evaluation: A Case Study in Meta-Evaluation for Long-Form QA Benchmarks
von: Hwang, Jena D., et al.
Veröffentlicht: (2026)
von: Hwang, Jena D., et al.
Veröffentlicht: (2026)
LitSearch: A Retrieval Benchmark for Scientific Literature Search
von: Ajith, Anirudh, et al.
Veröffentlicht: (2024)
von: Ajith, Anirudh, et al.
Veröffentlicht: (2024)
Neologism Learning for Controllability and Self-Verbalization
von: Hewitt, John, et al.
Veröffentlicht: (2025)
von: Hewitt, John, et al.
Veröffentlicht: (2025)
Demystifying Scientific Problem-Solving in LLMs by Probing Knowledge and Reasoning
von: Li, Alan, et al.
Veröffentlicht: (2025)
von: Li, Alan, et al.
Veröffentlicht: (2025)
Inferring Scientific Cross-Document Coreference and Hierarchy with Definition-Augmented Relational Reasoning
von: Forer, Lior, et al.
Veröffentlicht: (2024)
von: Forer, Lior, et al.
Veröffentlicht: (2024)
CHIMERA: A Knowledge Base of Scientific Idea Recombinations for Research Analysis and Ideation
von: Sternlicht, Noy, et al.
Veröffentlicht: (2025)
von: Sternlicht, Noy, et al.
Veröffentlicht: (2025)
Because we have LLMs, we Can and Should Pursue Agentic Interpretability
von: Kim, Been, et al.
Veröffentlicht: (2025)
von: Kim, Been, et al.
Veröffentlicht: (2025)
Answer, Assemble, Ace: Understanding How LMs Answer Multiple Choice Questions
von: Wiegreffe, Sarah, et al.
Veröffentlicht: (2024)
von: Wiegreffe, Sarah, et al.
Veröffentlicht: (2024)
Rural Education: Issues and Practice. Source Books on Education, Vol. 25. Garland Reference Library of Social Science, Vol. 473.
von: DeYoung, Alan J., Ed.
Veröffentlicht: (1991)
von: DeYoung, Alan J., Ed.
Veröffentlicht: (1991)
Understanding Usage and Engagement in AI-Powered Scientific Research Tools: The Asta Interaction Dataset
von: Haddad, Dany, et al.
Veröffentlicht: (2026)
von: Haddad, Dany, et al.
Veröffentlicht: (2026)
DISCOVERYWORLD: A Virtual Environment for Developing and Evaluating Automated Scientific Discovery Agents
von: Jansen, Peter, et al.
Veröffentlicht: (2024)
von: Jansen, Peter, et al.
Veröffentlicht: (2024)
Resource Sharing of Micro-Software, or, What Ever Happened to All That CP/M Compatibility?
von: DeYoung, Barbara
Veröffentlicht: (1984)
von: DeYoung, Barbara
Veröffentlicht: (1984)
SciRIFF: A Resource to Enhance Language Model Instruction-Following over Scientific Literature
von: Wadden, David, et al.
Veröffentlicht: (2024)
von: Wadden, David, et al.
Veröffentlicht: (2024)
OLMES: A Standard for Language Model Evaluations
von: Gu, Yuling, et al.
Veröffentlicht: (2024)
von: Gu, Yuling, et al.
Veröffentlicht: (2024)
AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite
von: Bragg, Jonathan, et al.
Veröffentlicht: (2025)
von: Bragg, Jonathan, et al.
Veröffentlicht: (2025)
Anagent For Enhancing Scientific Table & Figure Analysis
von: Guo, Xuehang, et al.
Veröffentlicht: (2026)
von: Guo, Xuehang, et al.
Veröffentlicht: (2026)
Downstream Trade-offs of a Family of Text Watermarks
von: Ajith, Anirudh, et al.
Veröffentlicht: (2023)
von: Ajith, Anirudh, et al.
Veröffentlicht: (2023)
Comparing Retrieval Strategies to Capture Interdisciplinary Scientific Research: A Bibliometric Evaluation of the Integration of Neuroscience and Computer Science
von: Isla, Malena Mendez, et al.
Veröffentlicht: (2025)
von: Isla, Malena Mendez, et al.
Veröffentlicht: (2025)
Semantic Structure of Feature Space in Large Language Models
von: Kozlowski, Austin C., et al.
Veröffentlicht: (2026)
von: Kozlowski, Austin C., et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
The (Short-Term) Effects of Large Language Models on Unemployment and Earnings
von: Chen, Danqing, et al.
Veröffentlicht: (2025) -
SciMON: Scientific Inspiration Machines Optimized for Novelty
von: Wang, Qingyun, et al.
Veröffentlicht: (2023) -
Generating Literature-Driven Scientific Theories at Scale
von: Jansen, Peter, et al.
Veröffentlicht: (2026) -
Polarization by Design: How Elites Could Shape Mass Preferences as AI Reduces Persuasion Costs
von: Kunievsky, Nadav
Veröffentlicht: (2025) -
Linear Regression in a Nonlinear World
von: Kunievsky, Nadav
Veröffentlicht: (2025)