BoxingGym: Benchmarking Progress in Automated Experimental Design and Model Discovery
Fuente:
arXiv
Saved in:
| Main Authors: | Gandhi, Kanishk, Li, Michael Y., Goodyear, Lyle, Bhatia, Agam, Li, Louise, Bhaskar, Aditi, Zaman, Mohammed, Goodman, Noah D. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning to Simulate Human Dialogue
by: Gandhi, Kanishk, et al.
Published: (2026)
by: Gandhi, Kanishk, et al.
Published: (2026)
Endless Terminals: Scaling RL Environments for Terminal Agents
by: Gandhi, Kanishk, et al.
Published: (2026)
by: Gandhi, Kanishk, et al.
Published: (2026)
Automated Statistical Model Discovery with Language Models
by: Li, Michael Y., et al.
Published: (2024)
by: Li, Michael Y., et al.
Published: (2024)
Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs
by: Gandhi, Kanishk, et al.
Published: (2025)
by: Gandhi, Kanishk, et al.
Published: (2025)
PRDP: Progressively Refined Differentiable Physics
by: Bhatia, Kanishk, et al.
Published: (2025)
by: Bhatia, Kanishk, et al.
Published: (2025)
Non-literal Understanding of Number Words by Language Models
by: Tsvilodub, Polina, et al.
Published: (2025)
by: Tsvilodub, Polina, et al.
Published: (2025)
Self-Supervised Alignment with Mutual Information: Learning to Follow Principles without Preference Labels
by: Fränken, Jan-Philipp, et al.
Published: (2024)
by: Fränken, Jan-Philipp, et al.
Published: (2024)
Procedural Dilemma Generation for Evaluating Moral Reasoning in Humans and Language Models
by: Fränken, Jan-Philipp, et al.
Published: (2024)
by: Fränken, Jan-Philipp, et al.
Published: (2024)
Scaling up the think-aloud method
by: Wurgaft, Daniel, et al.
Published: (2025)
by: Wurgaft, Daniel, et al.
Published: (2025)
Stream of Search (SoS): Learning to Search in Language
by: Gandhi, Kanishk, et al.
Published: (2024)
by: Gandhi, Kanishk, et al.
Published: (2024)
CriticAL: Critic Automation with Language Models
by: Li, Michael Y., et al.
Published: (2024)
by: Li, Michael Y., et al.
Published: (2024)
Psychometric Alignment: Capturing Human Knowledge Distributions via Language Models
by: He-Yueya, Joy, et al.
Published: (2024)
by: He-Yueya, Joy, et al.
Published: (2024)
AExGym: Benchmarks and Environments for Adaptive Experimentation
by: Wang, Jimmy, et al.
Published: (2024)
by: Wang, Jimmy, et al.
Published: (2024)
Reseña de "The Discursive Construction of European Identities: A Multi-Level Approach to Discourse and Identity in the Transforming European Union" de Michal Krzyzanowski.
by: Aditi Bhatia
Published: (2012)
by: Aditi Bhatia
Published: (2012)
Building community in international politics: A study of political press conferences
by: Aditi Bhatia
Published: (2011)
by: Aditi Bhatia
Published: (2011)
International genre, local flavour: Analysis of PetroChina’s Corporate and Social Responsibility Report
by: Aditi Bhatia
Published: (2013)
by: Aditi Bhatia
Published: (2013)
Automated Discovery of Tactic Libraries for Interactive Theorem Proving
by: Xin, Yutong, et al.
Published: (2025)
by: Xin, Yutong, et al.
Published: (2025)
The Effect of State Representation on LLM Agent Behavior in Dynamic Routing Games
by: Goodyear, Lyle, et al.
Published: (2025)
by: Goodyear, Lyle, et al.
Published: (2025)
ProgressGym: Alignment with a Millennium of Moral Progress
by: Qiu, Tianyi, et al.
Published: (2024)
by: Qiu, Tianyi, et al.
Published: (2024)
Human-like Affective Cognition in Foundation Models
by: Gandhi, Kanishk, et al.
Published: (2024)
by: Gandhi, Kanishk, et al.
Published: (2024)
PhysGym: Benchmarking LLMs in Interactive Physics Discovery with Controlled Priors
by: Chen, Yimeng, et al.
Published: (2025)
by: Chen, Yimeng, et al.
Published: (2025)
CrystalGym: A New Benchmark for Materials Discovery Using Reinforcement Learning
by: Govindarajan, Prashant, et al.
Published: (2025)
by: Govindarajan, Prashant, et al.
Published: (2025)
SeekerGym: A Benchmark for Reliable Information Seeking
by: Kim, Remy, et al.
Published: (2026)
by: Kim, Remy, et al.
Published: (2026)
A test the robustness of ASPIC based stock assessments using simulated Atlantic Blue Marlin data
by: Goodyear, C.P.
Published: (2000)
by: Goodyear, C.P.
Published: (2000)
ColorSwap: A Color and Word Order Dataset for Multimodal Evaluation
by: Burapacheep, Jirayu, et al.
Published: (2024)
by: Burapacheep, Jirayu, et al.
Published: (2024)
System Test Case Design from Requirements Specifications: Insights and Challenges of Using ChatGPT
by: Bhatia, Shreya, et al.
Published: (2024)
by: Bhatia, Shreya, et al.
Published: (2024)
Significance of Size-Dependent Hygroscopicity Parameterization for Aerosol-Cloud Interactions
by: Gohil, Kanishk
Published: (2025)
by: Gohil, Kanishk
Published: (2025)
Aiding Medical Diagnosis through Image Synthesis and Classification
by: Choudhary, Kanishk
Published: (2025)
by: Choudhary, Kanishk
Published: (2025)
The Inductive Bottleneck: Data-Driven Emergence of Representational Sparsity in Vision Transformers
by: Awadhiya, Kanishk
Published: (2025)
by: Awadhiya, Kanishk
Published: (2025)
Bifocal Attention: Harmonizing Geometric and Spectral Positional Embeddings for Algorithmic Generalization
by: Awadhiya, Kanishk
Published: (2026)
by: Awadhiya, Kanishk
Published: (2026)
D3-Gym: Constructing Real-World Verifiable Environments for Data-Driven Discovery
by: Moussa, Hanane Nour, et al.
Published: (2026)
by: Moussa, Hanane Nour, et al.
Published: (2026)
Is Child-Directed Speech Effective Training Data for Language Models?
by: Feng, Steven Y., et al.
Published: (2024)
by: Feng, Steven Y., et al.
Published: (2024)
DevOps-Gym: Benchmarking AI Agents in Software DevOps Cycle
by: Tang, Yuheng, et al.
Published: (2026)
by: Tang, Yuheng, et al.
Published: (2026)
Neural Garbage Collection: Learning to Forget while Learning to Reason
by: Li, Michael Y., et al.
Published: (2026)
by: Li, Michael Y., et al.
Published: (2026)
AMemGym: Interactive Memory Benchmarking for Assistants in Long-Horizon Conversations
by: Jiayang, Cheng, et al.
Published: (2026)
by: Jiayang, Cheng, et al.
Published: (2026)
Learning to Compress Prompts with Gist Tokens
by: Mu, Jesse, et al.
Published: (2023)
by: Mu, Jesse, et al.
Published: (2023)
Subliminal Learning Is Steering Vector Distillation
by: Blank, Camila, et al.
Published: (2026)
by: Blank, Camila, et al.
Published: (2026)
ArgCMV: An Argument Summarization Benchmark for the LLM-era
by: Gurjar, Omkar, et al.
Published: (2025)
by: Gurjar, Omkar, et al.
Published: (2025)
Automated Rewards via LLM-Generated Progress Functions
by: Sarukkai, Vishnu, et al.
Published: (2024)
by: Sarukkai, Vishnu, et al.
Published: (2024)
Exploring Topic Trends in COVID-19 Research Literature using Non-Negative Matrix Factorization
by: Patel, Divya, et al.
Published: (2025)
by: Patel, Divya, et al.
Published: (2025)
Similar Items
-
Learning to Simulate Human Dialogue
by: Gandhi, Kanishk, et al.
Published: (2026) -
Endless Terminals: Scaling RL Environments for Terminal Agents
by: Gandhi, Kanishk, et al.
Published: (2026) -
Automated Statistical Model Discovery with Language Models
by: Li, Michael Y., et al.
Published: (2024) -
Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs
by: Gandhi, Kanishk, et al.
Published: (2025) -
PRDP: Progressively Refined Differentiable Physics
by: Bhatia, Kanishk, et al.
Published: (2025)