PhysGym: Benchmarking LLMs in Interactive Physics Discovery with Controlled Priors
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Yimeng, Piȩkos, Piotr, Ostaszewski, Mateusz, Laakom, Firas, Schmidhuber, Jürgen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Huxley-Gödel Machine: Human-Level Coding Agent Development by an Approximation of the Optimal Self-Improving Machine
von: Wang, Wenyi, et al.
Veröffentlicht: (2025)
von: Wang, Wenyi, et al.
Veröffentlicht: (2025)
Fairness Overfitting in Machine Learning: An Information-Theoretic Perspective
von: Laakom, Firas, et al.
Veröffentlicht: (2025)
von: Laakom, Firas, et al.
Veröffentlicht: (2025)
FACTS: A Factored State-Space Framework For World Modelling
von: Nanbo, Li, et al.
Veröffentlicht: (2024)
von: Nanbo, Li, et al.
Veröffentlicht: (2024)
Planning to Explore: Curiosity-Driven Planning for LLM Test Generation
von: Amayuelas, Alfonso, et al.
Veröffentlicht: (2026)
von: Amayuelas, Alfonso, et al.
Veröffentlicht: (2026)
Hyperbolic Residual Quantization: Discrete Representations for Data with Latent Hierarchies
von: Piękos, Piotr, et al.
Veröffentlicht: (2025)
von: Piękos, Piotr, et al.
Veröffentlicht: (2025)
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing
von: Piękos, Piotr, et al.
Veröffentlicht: (2025)
von: Piękos, Piotr, et al.
Veröffentlicht: (2025)
EduGym: An Environment and Notebook Suite for Reinforcement Learning Education
von: Moerland, Thomas M., et al.
Veröffentlicht: (2023)
von: Moerland, Thomas M., et al.
Veröffentlicht: (2023)
BoxingGym: Benchmarking Progress in Automated Experimental Design and Model Discovery
von: Gandhi, Kanishk, et al.
Veröffentlicht: (2025)
von: Gandhi, Kanishk, et al.
Veröffentlicht: (2025)
Is Temporal Difference Learning the Gold Standard for Stitching in RL?
von: Bortkiewicz, Michał, et al.
Veröffentlicht: (2025)
von: Bortkiewicz, Michał, et al.
Veröffentlicht: (2025)
DM-Bench: Benchmarking LLMs for Personalized Decision Making in Diabetes Management
von: Cardei, Maria Ana, et al.
Veröffentlicht: (2025)
von: Cardei, Maria Ana, et al.
Veröffentlicht: (2025)
Cognitive Structure Generation: From Educational Priors to Policy Optimization
von: Gu, Hengnian, et al.
Veröffentlicht: (2025)
von: Gu, Hengnian, et al.
Veröffentlicht: (2025)
GEM: A Gym for Agentic LLMs
von: Liu, Zichen, et al.
Veröffentlicht: (2025)
von: Liu, Zichen, et al.
Veröffentlicht: (2025)
SwitchHead: Accelerating Transformers with Mixture-of-Experts Attention
von: Csordás, Róbert, et al.
Veröffentlicht: (2023)
von: Csordás, Róbert, et al.
Veröffentlicht: (2023)
Deep Learning Opacity in Scientific Discovery
von: Duede, Eamon
Veröffentlicht: (2022)
von: Duede, Eamon
Veröffentlicht: (2022)
D3-Gym: Constructing Real-World Verifiable Environments for Data-Driven Discovery
von: Moussa, Hanane Nour, et al.
Veröffentlicht: (2026)
von: Moussa, Hanane Nour, et al.
Veröffentlicht: (2026)
ProgressGym: Alignment with a Millennium of Moral Progress
von: Qiu, Tianyi, et al.
Veröffentlicht: (2024)
von: Qiu, Tianyi, et al.
Veröffentlicht: (2024)
Benchmarking the Legal Reasoning of LLMs in Arabic Islamic Inheritance Cases
von: AlDahoul, Nouar, et al.
Veröffentlicht: (2025)
von: AlDahoul, Nouar, et al.
Veröffentlicht: (2025)
Prerequisite Structure Discovery in Intelligent Tutoring Systems
von: Annabi, Louis, et al.
Veröffentlicht: (2024)
von: Annabi, Louis, et al.
Veröffentlicht: (2024)
Diffusion-Inspired Cold Start with Sufficient Prior in Computerized Adaptive Testing
von: Ma, Haiping, et al.
Veröffentlicht: (2024)
von: Ma, Haiping, et al.
Veröffentlicht: (2024)
Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components
von: Potham, Ram
Veröffentlicht: (2025)
von: Potham, Ram
Veröffentlicht: (2025)
PersonaGym: Evaluating Persona Agents and LLMs
von: Samuel, Vinay, et al.
Veröffentlicht: (2024)
von: Samuel, Vinay, et al.
Veröffentlicht: (2024)
Say My Name: a Model's Bias Discovery Framework
von: Ciranni, Massimiliano, et al.
Veröffentlicht: (2024)
von: Ciranni, Massimiliano, et al.
Veröffentlicht: (2024)
Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple Interactions
von: Chan, Yik Siu, et al.
Veröffentlicht: (2025)
von: Chan, Yik Siu, et al.
Veröffentlicht: (2025)
Leveraging Large Language Models for Tacit Knowledge Discovery in Organizational Contexts
von: Zuin, Gianlucca, et al.
Veröffentlicht: (2025)
von: Zuin, Gianlucca, et al.
Veröffentlicht: (2025)
NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security
von: Shao, Minghao, et al.
Veröffentlicht: (2024)
von: Shao, Minghao, et al.
Veröffentlicht: (2024)
Interestingness as an Inductive Heuristic for Future Compression Progress
von: Herrmann, Vincent, et al.
Veröffentlicht: (2026)
von: Herrmann, Vincent, et al.
Veröffentlicht: (2026)
FFB: A Fair Fairness Benchmark for In-Processing Group Fairness Methods
von: Han, Xiaotian, et al.
Veröffentlicht: (2023)
von: Han, Xiaotian, et al.
Veröffentlicht: (2023)
Deprecating Benchmarks: Criteria and Framework
von: Joaquin, Ayrton San, et al.
Veröffentlicht: (2025)
von: Joaquin, Ayrton San, et al.
Veröffentlicht: (2025)
Defining and Evaluating Physical Safety for Large Language Models
von: Tang, Yung-Chen, et al.
Veröffentlicht: (2024)
von: Tang, Yung-Chen, et al.
Veröffentlicht: (2024)
MMAI Gym for Science: Training Liquid Foundation Models for Drug Discovery
von: Kuznetsov, Maksim, et al.
Veröffentlicht: (2026)
von: Kuznetsov, Maksim, et al.
Veröffentlicht: (2026)
UserBench: An Interactive Gym Environment for User-Centric Agents
von: Qian, Cheng, et al.
Veröffentlicht: (2025)
von: Qian, Cheng, et al.
Veröffentlicht: (2025)
Gym4ReaL: A Suite for Benchmarking Real-World Reinforcement Learning
von: Salaorni, Davide, et al.
Veröffentlicht: (2025)
von: Salaorni, Davide, et al.
Veröffentlicht: (2025)
Healthcare LLM Benchmarks Are Only as Good as Their Explicit Assumptions
von: Raman, Naveen, et al.
Veröffentlicht: (2026)
von: Raman, Naveen, et al.
Veröffentlicht: (2026)
vTune: Verifiable Fine-Tuning for LLMs Through Backdooring
von: Zhang, Eva, et al.
Veröffentlicht: (2024)
von: Zhang, Eva, et al.
Veröffentlicht: (2024)
HeuriGym: An Agentic Benchmark for LLM-Crafted Heuristics in Combinatorial Optimization
von: Chen, Hongzheng, et al.
Veröffentlicht: (2025)
von: Chen, Hongzheng, et al.
Veröffentlicht: (2025)
MateInfoUB: A Real-World Benchmark for Testing LLMs in Competitive, Multilingual, and Multimodal Educational Tasks
von: Marius, Dumitran Adrian, et al.
Veröffentlicht: (2025)
von: Marius, Dumitran Adrian, et al.
Veröffentlicht: (2025)
GECOBench: A Gender-Controlled Text Dataset and Benchmark for Quantifying Biases in Explanations
von: Wilming, Rick, et al.
Veröffentlicht: (2024)
von: Wilming, Rick, et al.
Veröffentlicht: (2024)
Quantifying Prediction Consistency Under Fine-Tuning Multiplicity in Tabular LLMs
von: Hamman, Faisal, et al.
Veröffentlicht: (2024)
von: Hamman, Faisal, et al.
Veröffentlicht: (2024)
MoralityGym: A Benchmark for Evaluating Hierarchical Moral Alignment in Sequential Decision-Making Agents
von: Rosen, Simon, et al.
Veröffentlicht: (2026)
von: Rosen, Simon, et al.
Veröffentlicht: (2026)
TimeSeriesGym: A Scalable Benchmark for (Time Series) Machine Learning Engineering Agents
von: Cai, Yifu, et al.
Veröffentlicht: (2025)
von: Cai, Yifu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Huxley-Gödel Machine: Human-Level Coding Agent Development by an Approximation of the Optimal Self-Improving Machine
von: Wang, Wenyi, et al.
Veröffentlicht: (2025) -
Fairness Overfitting in Machine Learning: An Information-Theoretic Perspective
von: Laakom, Firas, et al.
Veröffentlicht: (2025) -
FACTS: A Factored State-Space Framework For World Modelling
von: Nanbo, Li, et al.
Veröffentlicht: (2024) -
Planning to Explore: Curiosity-Driven Planning for LLM Test Generation
von: Amayuelas, Alfonso, et al.
Veröffentlicht: (2026) -
Hyperbolic Residual Quantization: Discrete Representations for Data with Latent Hierarchies
von: Piękos, Piotr, et al.
Veröffentlicht: (2025)