DiscoveryBench: Towards Data-Driven Discovery with Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Majumder, Bodhisattwa Prasad, Surana, Harshit, Agarwal, Dhruv, Mishra, Bhavana Dalvi, Meena, Abhijeetsingh, Prakhar, Aryan, Vora, Tirth, Khot, Tushar, Sabharwal, Ashish, Clark, Peter |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Data-driven Discovery with Large Generative Models
by: Majumder, Bodhisattwa Prasad, et al.
Published: (2024)
by: Majumder, Bodhisattwa Prasad, et al.
Published: (2024)
Latent Factor Models Meets Instructions: Goal-conditioned Latent Factor Discovery without Task Supervision
by: Xie, Zhouhang, et al.
Published: (2025)
by: Xie, Zhouhang, et al.
Published: (2025)
AutoDiscovery: Open-ended Scientific Discovery via Bayesian Surprise
by: Agarwal, Dhruv, et al.
Published: (2025)
by: Agarwal, Dhruv, et al.
Published: (2025)
DISCOVERYWORLD: A Virtual Environment for Developing and Evaluating Automated Scientific Discovery Agents
by: Jansen, Peter, et al.
Published: (2024)
by: Jansen, Peter, et al.
Published: (2024)
CodeScientist: End-to-End Semi-Automated Scientific Discovery with Code-based Experimentation
by: Jansen, Peter, et al.
Published: (2025)
by: Jansen, Peter, et al.
Published: (2025)
Skill Set Optimization: Reinforcing Language Model Behavior via Transferable Skills
by: Nottingham, Kolby, et al.
Published: (2024)
by: Nottingham, Kolby, et al.
Published: (2024)
ArtifactLinker: Linking Scientific Artifacts for Automatic State-of-the-Art Discovery
by: Yu, Haofei, et al.
Published: (2026)
by: Yu, Haofei, et al.
Published: (2026)
ADaPT: As-Needed Decomposition and Planning with Language Models
by: Prasad, Archiki, et al.
Published: (2023)
by: Prasad, Archiki, et al.
Published: (2023)
BaRDa: A Belief and Reasoning Dataset that Separates Factual Accuracy and Reasoning Ability
by: Clark, Peter, et al.
Published: (2023)
by: Clark, Peter, et al.
Published: (2023)
Leveraging In-Context Learning for Language Model Agents
by: Gupta, Shivanshu, et al.
Published: (2025)
by: Gupta, Shivanshu, et al.
Published: (2025)
To Tell The Truth: Language of Deception and Language Models
by: Hazra, Sanchaita, et al.
Published: (2023)
by: Hazra, Sanchaita, et al.
Published: (2023)
Bias Runs Deep: Implicit Reasoning Biases in Persona-Assigned LLMs
by: Gupta, Shashank, et al.
Published: (2023)
by: Gupta, Shashank, et al.
Published: (2023)
VoiceAgentBench: Are Voice Assistants ready for agentic tasks?
by: Jain, Dhruv, et al.
Published: (2025)
by: Jain, Dhruv, et al.
Published: (2025)
SUPER: Evaluating Agents on Setting Up and Executing Tasks from Research Repositories
by: Bogin, Ben, et al.
Published: (2024)
by: Bogin, Ben, et al.
Published: (2024)
Tell, Don't Show!: Language Guidance Eases Transfer Across Domains in Images and Videos
by: Kalluri, Tarun, et al.
Published: (2024)
by: Kalluri, Tarun, et al.
Published: (2024)
AI Safety Should Prioritize the Future of Work
by: Hazra, Sanchaita, et al.
Published: (2025)
by: Hazra, Sanchaita, et al.
Published: (2025)
HARPA: A Testability-Driven, Literature-Grounded Framework for Research Ideation
by: Vasu, Rosni, et al.
Published: (2025)
by: Vasu, Rosni, et al.
Published: (2025)
Lightweight Query Routing for Adaptive RAG: A Baseline Study on RAGRouter-Bench
by: Bansal, Prakhar, et al.
Published: (2026)
by: Bansal, Prakhar, et al.
Published: (2026)
DESIGN TOKENS AND CONTRACT-FIRST SYSTEMS FOR CROSS-PLATFORM UI CONSISTENCY
by: Harshit Sunilkumar Vora
Published: (2026)
by: Harshit Sunilkumar Vora
Published: (2026)
MoNaCo: More Natural and Complex Questions for Reasoning Across Dozens of Documents
by: Wolfson, Tomer, et al.
Published: (2025)
by: Wolfson, Tomer, et al.
Published: (2025)
Accepted with Minor Revisions: Value of AI-Assisted Scientific Writing
by: Hazra, Sanchaita, et al.
Published: (2025)
by: Hazra, Sanchaita, et al.
Published: (2025)
The Good, the Bad, and the Ugly: The Role of AI Quality Disclosure in Lie Detection
by: Bhattacharya, Haimanti, et al.
Published: (2024)
by: Bhattacharya, Haimanti, et al.
Published: (2024)
HypER: Literature-grounded Hypothesis Generation and Distillation with Provenance
by: Vasu, Rosni, et al.
Published: (2025)
by: Vasu, Rosni, et al.
Published: (2025)
AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite
by: Bragg, Jonathan, et al.
Published: (2025)
by: Bragg, Jonathan, et al.
Published: (2025)
Tailoring with Targeted Precision: Edit-Based Agents for Open-Domain Procedure Customization
by: Lal, Yash Kumar, et al.
Published: (2023)
by: Lal, Yash Kumar, et al.
Published: (2023)
AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents
by: Trivedi, Harsh, et al.
Published: (2024)
by: Trivedi, Harsh, et al.
Published: (2024)
Scaling up Discovery of Latent Concepts in Deep NLP Models
by: Hawasly, Majd, et al.
Published: (2023)
by: Hawasly, Majd, et al.
Published: (2023)
Put Your Money Where Your Mouth Is: Evaluating Strategic Planning and Execution of LLM Agents in an Auction Arena
by: Chen, Jiangjie, et al.
Published: (2023)
by: Chen, Jiangjie, et al.
Published: (2023)
Leveraging Code to Improve In-context Learning for Semantic Parsing
by: Bogin, Ben, et al.
Published: (2023)
by: Bogin, Ben, et al.
Published: (2023)
A Little Depth Goes a Long Way: The Expressive Power of Log-Depth Transformers
by: Merrill, William, et al.
Published: (2025)
by: Merrill, William, et al.
Published: (2025)
Exact Expressive Power of Transformers with Padding
by: Merrill, William, et al.
Published: (2025)
by: Merrill, William, et al.
Published: (2025)
The Expressive Power of Transformers with Chain of Thought
by: Merrill, William, et al.
Published: (2023)
by: Merrill, William, et al.
Published: (2023)
A Logic for Expressing Log-Precision Transformers
by: Merrill, William, et al.
Published: (2022)
by: Merrill, William, et al.
Published: (2022)
On the Reasoning Abilities of Masked Diffusion Language Models
by: Svete, Anej, et al.
Published: (2025)
by: Svete, Anej, et al.
Published: (2025)
Dynamic Sensor Selection for Biomarker Discovery
by: Pickard, Joshua, et al.
Published: (2024)
by: Pickard, Joshua, et al.
Published: (2024)
AI as a Medical Ally: Evaluating ChatGPT's Usage and Impact in Indian Healthcare
by: Raina, Aryaman, et al.
Published: (2024)
by: Raina, Aryaman, et al.
Published: (2024)
Beyond the Parameters: A Technical Survey of Contextual Enrichment in Large Language Models: From In-Context Prompting to Causal Retrieval-Augmented Generation
by: Bansal, Prakhar, et al.
Published: (2026)
by: Bansal, Prakhar, et al.
Published: (2026)
Participation in First Proof Challenge by Zetesis Labs leveraging our work on formal Epistemology for Proof Discovery
by: Gupta, Dhruv
Published: (2026)
by: Gupta, Dhruv
Published: (2026)
Auto-Discovery-Bench: Diagnosing Structured State Tracking in Oracle-Guided Discovery
by: Chen, Tingting, et al.
Published: (2025)
by: Chen, Tingting, et al.
Published: (2025)
Sequential Causal Discovery with Noisy Language Model Priors
by: Verma, Prakhar, et al.
Published: (2025)
by: Verma, Prakhar, et al.
Published: (2025)
Similar Items
-
Data-driven Discovery with Large Generative Models
by: Majumder, Bodhisattwa Prasad, et al.
Published: (2024) -
Latent Factor Models Meets Instructions: Goal-conditioned Latent Factor Discovery without Task Supervision
by: Xie, Zhouhang, et al.
Published: (2025) -
AutoDiscovery: Open-ended Scientific Discovery via Bayesian Surprise
by: Agarwal, Dhruv, et al.
Published: (2025) -
DISCOVERYWORLD: A Virtual Environment for Developing and Evaluating Automated Scientific Discovery Agents
by: Jansen, Peter, et al.
Published: (2024) -
CodeScientist: End-to-End Semi-Automated Scientific Discovery with Code-based Experimentation
by: Jansen, Peter, et al.
Published: (2025)