Self-Driving Datasets: From 20 Million Papers to Nuanced Biomedical Knowledge at Scale
Fuente:
arXiv
Saved in:
| Main Authors: | Jones, Haydn, Zeng, Yimeng, Rose, Alden, Yifei, Li S., Huang, Yining, Wu, Kaiwen, Liang, Jiaming, Huan, Maggie Ziyu, Barash, Yoseph, de la Fuente-Nunez, Cesar, Bastani, Osbert, Ives, Zachary, Yatskar, Mark, Gardner, Jacob R. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Dataset for Distilling Knowledge Priors from Literature for Therapeutic Design
by: Jones, Haydn Thomas, et al.
Published: (2025)
by: Jones, Haydn Thomas, et al.
Published: (2025)
Improving Structural Diversity of Blackbox LLMs via Chain-of-Specification Prompting
by: Young, Halley, et al.
Published: (2024)
by: Young, Halley, et al.
Published: (2024)
Purely Agent-Driven Black-Box Optimization for Biological Design
by: Maus, Natalie, et al.
Published: (2026)
by: Maus, Natalie, et al.
Published: (2026)
Adversarial Query Synthesis via Bayesian Optimization
by: Tao, Jeffrey, et al.
Published: (2026)
by: Tao, Jeffrey, et al.
Published: (2026)
Asymptotic Normality of Generalized Low-Rank Matrix Sensing via Riemannian Geometry
by: Bastani, Osbert
Published: (2024)
by: Bastani, Osbert
Published: (2024)
Generative Adversarial Model-Based Optimization via Source Critic Regularization
by: Yao, Michael S., et al.
Published: (2024)
by: Yao, Michael S., et al.
Published: (2024)
Winner's Curse Drives False Promises in Data-Driven Decisions: A Case Study in Refugee Matching
by: Bastani, Hamsa, et al.
Published: (2026)
by: Bastani, Hamsa, et al.
Published: (2026)
Stochastic Online Conformal Prediction with Semi-Bandit Feedback
by: Ge, Haosen, et al.
Published: (2024)
by: Ge, Haosen, et al.
Published: (2024)
Are AI Capabilities Increasing Exponentially? A Competing Hypothesis
by: Ge, Haosen, et al.
Published: (2026)
by: Ge, Haosen, et al.
Published: (2026)
Rethinking Algorithmic Fairness for Human-AI Collaboration
by: Ge, Haosen, et al.
Published: (2023)
by: Ge, Haosen, et al.
Published: (2023)
Beating the Winner's Curse via Inference-Aware Policy Optimization
by: Bastani, Hamsa, et al.
Published: (2025)
by: Bastani, Hamsa, et al.
Published: (2025)
Improving Human Sequential Decision-Making with Reinforcement Learning
by: Bastani, Hamsa, et al.
Published: (2021)
by: Bastani, Hamsa, et al.
Published: (2021)
Large Scale Multi-Task Bayesian Optimization with Large Language Models
by: Zeng, Yimeng, et al.
Published: (2025)
by: Zeng, Yimeng, et al.
Published: (2025)
Synthesizing Trajectory Queries from Examples
by: Mell, Stephen, et al.
Published: (2026)
by: Mell, Stephen, et al.
Published: (2026)
Group-Sparse Matrix Factorization for Transfer Learning of Word Embeddings
by: Xu, Kan, et al.
Published: (2021)
by: Xu, Kan, et al.
Published: (2021)
Stochastic Bandits with ReLU Neural Networks
by: Xu, Kan, et al.
Published: (2024)
by: Xu, Kan, et al.
Published: (2024)
Policy-Guided Stepwise Model Routing for Cost-Effective Reasoning
by: Si, Wenwen, et al.
Published: (2026)
by: Si, Wenwen, et al.
Published: (2026)
Decaf: Improving Neural Decompilation with Automatic Feedback and Search
by: Shypula, Alexander, et al.
Published: (2026)
by: Shypula, Alexander, et al.
Published: (2026)
Conformal Structured Prediction
by: Zhang, Botong, et al.
Published: (2024)
by: Zhang, Botong, et al.
Published: (2024)
Uncertainty Quantification for Neurosymbolic Programs via Compositional Conformal Prediction
by: Ramalingam, Ramya, et al.
Published: (2024)
by: Ramalingam, Ramya, et al.
Published: (2024)
LLM Program Optimization via Retrieval Augmented Search
by: Anupam, Sagnik, et al.
Published: (2025)
by: Anupam, Sagnik, et al.
Published: (2025)
Prior-Agnostic Incentive-Compatible Exploration
by: Ramalingam, Ramya, et al.
Published: (2026)
by: Ramalingam, Ramya, et al.
Published: (2026)
Optimal Program Synthesis via Abstract Interpretation
by: Mell, Stephen, et al.
Published: (2026)
by: Mell, Stephen, et al.
Published: (2026)
Diversity By Design: Leveraging Distribution Matching for Offline Model-Based Optimization
by: Yao, Michael S., et al.
Published: (2025)
by: Yao, Michael S., et al.
Published: (2025)
SPARLING: Learning Latent Representations with Extremely Sparse Activations
by: Gupta, Kavi, et al.
Published: (2023)
by: Gupta, Kavi, et al.
Published: (2023)
Conformal Constrained Policy Optimization for Cost-Effective LLM Agents
by: Si, Wenwen, et al.
Published: (2025)
by: Si, Wenwen, et al.
Published: (2025)
TRAQ: Trustworthy Retrieval Augmented Question Answering via Conformal Prediction
by: Li, Shuo, et al.
Published: (2023)
by: Li, Shuo, et al.
Published: (2023)
Choose, Don't Label: Multiple-Choice Query Synthesis for Program Disambiguation
by: Barnaby, Celeste, et al.
Published: (2026)
by: Barnaby, Celeste, et al.
Published: (2026)
Opportunistically Parallel Lambda Calculus
by: Mell, Stephen, et al.
Published: (2024)
by: Mell, Stephen, et al.
Published: (2024)
SeekerGym: A Benchmark for Reliable Information Seeking
by: Kim, Remy, et al.
Published: (2026)
by: Kim, Remy, et al.
Published: (2026)
Learning Performance-Improving Code Edits
by: Shypula, Alexander, et al.
Published: (2023)
by: Shypula, Alexander, et al.
Published: (2023)
Large-Scale Gaussian Processes via Alternating Projection
by: Wu, Kaiwen, et al.
Published: (2023)
by: Wu, Kaiwen, et al.
Published: (2023)
Effective Reinforcement Learning for Reasoning in Language Models
by: Huang, Lianghuan, et al.
Published: (2025)
by: Huang, Lianghuan, et al.
Published: (2025)
PopPy: Opportunistically Exploiting Parallelism in Python Compound AI Applications
by: Mell, Stephen, et al.
Published: (2026)
by: Mell, Stephen, et al.
Published: (2026)
RAPID: An Efficient Reinforcement Learning Algorithm for Small Language Models
by: Huang, Lianghuan, et al.
Published: (2025)
by: Huang, Lianghuan, et al.
Published: (2025)
Active Learning for Neurosymbolic Program Synthesis
by: Barnaby, Celeste, et al.
Published: (2025)
by: Barnaby, Celeste, et al.
Published: (2025)
Learned Offline Query Planning via Bayesian Optimization
by: Tao, Jeffrey, et al.
Published: (2025)
by: Tao, Jeffrey, et al.
Published: (2025)
ResearchQA: Evaluating Scholarly Question Answering at Scale Across 75 Fields with Survey-Mined Questions and Rubrics
by: Yifei, Li S., et al.
Published: (2025)
by: Yifei, Li S., et al.
Published: (2025)
BrowserArena: Evaluating LLM Agents on Real-World Web Navigation Tasks
by: Anupam, Sagnik, et al.
Published: (2025)
by: Anupam, Sagnik, et al.
Published: (2025)
One-Shot Safety Alignment for Large Language Models via Optimal Dualization
by: Huang, Xinmeng, et al.
Published: (2024)
by: Huang, Xinmeng, et al.
Published: (2024)
Similar Items
-
A Dataset for Distilling Knowledge Priors from Literature for Therapeutic Design
by: Jones, Haydn Thomas, et al.
Published: (2025) -
Improving Structural Diversity of Blackbox LLMs via Chain-of-Specification Prompting
by: Young, Halley, et al.
Published: (2024) -
Purely Agent-Driven Black-Box Optimization for Biological Design
by: Maus, Natalie, et al.
Published: (2026) -
Adversarial Query Synthesis via Bayesian Optimization
by: Tao, Jeffrey, et al.
Published: (2026) -
Asymptotic Normality of Generalized Low-Rank Matrix Sensing via Riemannian Geometry
by: Bastani, Osbert
Published: (2024)