DISCOVERYWORLD: A Virtual Environment for Developing and Evaluating Automated Scientific Discovery Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Jansen, Peter, Côté, Marc-Alexandre, Khot, Tushar, Bransom, Erin, Mishra, Bhavana Dalvi, Majumder, Bodhisattwa Prasad, Tafjord, Oyvind, Clark, Peter |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CodeScientist: End-to-End Semi-Automated Scientific Discovery with Code-based Experimentation
by: Jansen, Peter, et al.
Published: (2025)
by: Jansen, Peter, et al.
Published: (2025)
BaRDa: A Belief and Reasoning Dataset that Separates Factual Accuracy and Reasoning Ability
by: Clark, Peter, et al.
Published: (2023)
by: Clark, Peter, et al.
Published: (2023)
Latent Factor Models Meets Instructions: Goal-conditioned Latent Factor Discovery without Task Supervision
by: Xie, Zhouhang, et al.
Published: (2025)
by: Xie, Zhouhang, et al.
Published: (2025)
DiscoveryBench: Towards Data-Driven Discovery with Large Language Models
by: Majumder, Bodhisattwa Prasad, et al.
Published: (2024)
by: Majumder, Bodhisattwa Prasad, et al.
Published: (2024)
Skill Set Optimization: Reinforcing Language Model Behavior via Transferable Skills
by: Nottingham, Kolby, et al.
Published: (2024)
by: Nottingham, Kolby, et al.
Published: (2024)
AutoDiscovery: Open-ended Scientific Discovery via Bayesian Surprise
by: Agarwal, Dhruv, et al.
Published: (2025)
by: Agarwal, Dhruv, et al.
Published: (2025)
ArtifactLinker: Linking Scientific Artifacts for Automatic State-of-the-Art Discovery
by: Yu, Haofei, et al.
Published: (2026)
by: Yu, Haofei, et al.
Published: (2026)
From Models to Microtheories: Distilling a Model's Topical Knowledge for Grounded Question Answering
by: Weir, Nathaniel, et al.
Published: (2024)
by: Weir, Nathaniel, et al.
Published: (2024)
SUPER: Evaluating Agents on Setting Up and Executing Tasks from Research Repositories
by: Bogin, Ben, et al.
Published: (2024)
by: Bogin, Ben, et al.
Published: (2024)
Digital Socrates: Evaluating LLMs through Explanation Critiques
by: Gu, Yuling, et al.
Published: (2023)
by: Gu, Yuling, et al.
Published: (2023)
Enhancing Systematic Decompositional Natural Language Inference Using Informal Logic
by: Weir, Nathaniel, et al.
Published: (2024)
by: Weir, Nathaniel, et al.
Published: (2024)
Data-driven Discovery with Large Generative Models
by: Majumder, Bodhisattwa Prasad, et al.
Published: (2024)
by: Majumder, Bodhisattwa Prasad, et al.
Published: (2024)
HARPA: A Testability-Driven, Literature-Grounded Framework for Research Ideation
by: Vasu, Rosni, et al.
Published: (2025)
by: Vasu, Rosni, et al.
Published: (2025)
To Tell The Truth: Language of Deception and Language Models
by: Hazra, Sanchaita, et al.
Published: (2023)
by: Hazra, Sanchaita, et al.
Published: (2023)
Accepted with Minor Revisions: Value of AI-Assisted Scientific Writing
by: Hazra, Sanchaita, et al.
Published: (2025)
by: Hazra, Sanchaita, et al.
Published: (2025)
Tailoring with Targeted Precision: Edit-Based Agents for Open-Domain Procedure Customization
by: Lal, Yash Kumar, et al.
Published: (2023)
by: Lal, Yash Kumar, et al.
Published: (2023)
HypER: Literature-grounded Hypothesis Generation and Distillation with Provenance
by: Vasu, Rosni, et al.
Published: (2025)
by: Vasu, Rosni, et al.
Published: (2025)
Tell, Don't Show!: Language Guidance Eases Transfer Across Domains in Images and Videos
by: Kalluri, Tarun, et al.
Published: (2024)
by: Kalluri, Tarun, et al.
Published: (2024)
AI Safety Should Prioritize the Future of Work
by: Hazra, Sanchaita, et al.
Published: (2025)
by: Hazra, Sanchaita, et al.
Published: (2025)
Put Your Money Where Your Mouth Is: Evaluating Strategic Planning and Execution of LLM Agents in an Auction Arena
by: Chen, Jiangjie, et al.
Published: (2023)
by: Chen, Jiangjie, et al.
Published: (2023)
ADaPT: As-Needed Decomposition and Planning with Language Models
by: Prasad, Archiki, et al.
Published: (2023)
by: Prasad, Archiki, et al.
Published: (2023)
The Good, the Bad, and the Ugly: The Role of AI Quality Disclosure in Lie Detection
by: Bhattacharya, Haimanti, et al.
Published: (2024)
by: Bhattacharya, Haimanti, et al.
Published: (2024)
The Scientific Contribution Graph: Automated Literature-based Technological Roadmapping at Scale
by: Jansen, Peter A.
Published: (2026)
by: Jansen, Peter A.
Published: (2026)
Can Language Models Serve as Text-Based World Simulators?
by: Wang, Ruoyao, et al.
Published: (2024)
by: Wang, Ruoyao, et al.
Published: (2024)
SimpleToM: Exposing the Gap between Explicit ToM Inference and Implicit ToM Application in LLMs
by: Gu, Yuling, et al.
Published: (2024)
by: Gu, Yuling, et al.
Published: (2024)
Generating Literature-Driven Scientific Theories at Scale
by: Jansen, Peter, et al.
Published: (2026)
by: Jansen, Peter, et al.
Published: (2026)
Neologism Learning for Controllability and Self-Verbalization
by: Hewitt, John, et al.
Published: (2025)
by: Hewitt, John, et al.
Published: (2025)
AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite
by: Bragg, Jonathan, et al.
Published: (2025)
by: Bragg, Jonathan, et al.
Published: (2025)
Husky: A Unified, Open-Source Language Agent for Multi-Step Reasoning
by: Kim, Joongwon, et al.
Published: (2024)
by: Kim, Joongwon, et al.
Published: (2024)
Bias Runs Deep: Implicit Reasoning Biases in Persona-Assigned LLMs
by: Gupta, Shashank, et al.
Published: (2023)
by: Gupta, Shashank, et al.
Published: (2023)
ARIES: A Corpus of Scientific Paper Edits Made in Response to Peer Reviews
by: D'Arcy, Mike, et al.
Published: (2023)
by: D'Arcy, Mike, et al.
Published: (2023)
Leveraging In-Context Learning for Language Model Agents
by: Gupta, Shivanshu, et al.
Published: (2025)
by: Gupta, Shivanshu, et al.
Published: (2025)
Because we have LLMs, we Can and Should Pursue Agentic Interpretability
by: Kim, Been, et al.
Published: (2025)
by: Kim, Been, et al.
Published: (2025)
Answer, Assemble, Ace: Understanding How LMs Answer Multiple Choice Questions
by: Wiegreffe, Sarah, et al.
Published: (2024)
by: Wiegreffe, Sarah, et al.
Published: (2024)
Toward Automated Scientific Discovery in Hydrology: The Opportunities and Dangers of AI Augmented Research Frameworks
by: Darri Eythorsson, et al.
Published: (2025)
by: Darri Eythorsson, et al.
Published: (2025)
CodeDistiller: Automatically Generating Code Libraries for Scientific Coding Agents
by: Jansen, Peter, et al.
Published: (2025)
by: Jansen, Peter, et al.
Published: (2025)
CHIME: LLM-Assisted Hierarchical Organization of Scientific Studies for Literature Review Support
by: Hsu, Chao-Chun, et al.
Published: (2024)
by: Hsu, Chao-Chun, et al.
Published: (2024)
CARE: Extracting Experimental Findings From Clinical Literature
by: Naik, Aakanksha, et al.
Published: (2023)
by: Naik, Aakanksha, et al.
Published: (2023)
PreScience: A Benchmark for Forecasting Scientific Contributions
by: Ajith, Anirudh, et al.
Published: (2026)
by: Ajith, Anirudh, et al.
Published: (2026)
OLMES: A Standard for Language Model Evaluations
by: Gu, Yuling, et al.
Published: (2024)
by: Gu, Yuling, et al.
Published: (2024)
Similar Items
-
CodeScientist: End-to-End Semi-Automated Scientific Discovery with Code-based Experimentation
by: Jansen, Peter, et al.
Published: (2025) -
BaRDa: A Belief and Reasoning Dataset that Separates Factual Accuracy and Reasoning Ability
by: Clark, Peter, et al.
Published: (2023) -
Latent Factor Models Meets Instructions: Goal-conditioned Latent Factor Discovery without Task Supervision
by: Xie, Zhouhang, et al.
Published: (2025) -
DiscoveryBench: Towards Data-Driven Discovery with Large Language Models
by: Majumder, Bodhisattwa Prasad, et al.
Published: (2024) -
Skill Set Optimization: Reinforcing Language Model Behavior via Transferable Skills
by: Nottingham, Kolby, et al.
Published: (2024)