GSM-Agent: Understanding Agentic Reasoning Using Controllable Environments
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Hanlin, Guo, Tianyu, Mei, Song, Russell, Stuart, Ghosh, Nikhil, Bietti, Alberto, Jiao, Jiantao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Understanding the Mechanisms of Fast Hyperparameter Transfer
by: Ghosh, Nikhil, et al.
Published: (2025)
by: Ghosh, Nikhil, et al.
Published: (2025)
How Do LLMs Perform Two-Hop Reasoning in Context?
by: Guo, Tianyu, et al.
Published: (2025)
by: Guo, Tianyu, et al.
Published: (2025)
Generalization or Hallucination? Understanding Out-of-Context Reasoning in Transformers
by: Huang, Yixiao, et al.
Published: (2025)
by: Huang, Yixiao, et al.
Published: (2025)
Avoiding Catastrophe in Online Learning by Asking for Help
by: Plaut, Benjamin, et al.
Published: (2024)
by: Plaut, Benjamin, et al.
Published: (2024)
GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models
by: Mirzadeh, Iman, et al.
Published: (2024)
by: Mirzadeh, Iman, et al.
Published: (2024)
Safe Learning Under Irreversible Dynamics via Asking for Help
by: Plaut, Benjamin, et al.
Published: (2025)
by: Plaut, Benjamin, et al.
Published: (2025)
Token Assorted: Mixing Latent and Text Tokens for Improved Language Model Reasoning
by: Su, DiJia, et al.
Published: (2025)
by: Su, DiJia, et al.
Published: (2025)
Distributional Associations vs In-Context Reasoning: A Study of Feed-forward and Attention Layers
by: Chen, Lei, et al.
Published: (2024)
by: Chen, Lei, et al.
Published: (2024)
The Effective Horizon Explains Deep RL Performance in Stochastic Environments
by: Laidlaw, Cassidy, et al.
Published: (2023)
by: Laidlaw, Cassidy, et al.
Published: (2023)
Reasoning by Superposition: A Theoretical Perspective on Chain of Continuous Thought
by: Zhu, Hanlin, et al.
Published: (2025)
by: Zhu, Hanlin, et al.
Published: (2025)
Learning Associative Memories with Gradient Descent
by: Cabannes, Vivien, et al.
Published: (2024)
by: Cabannes, Vivien, et al.
Published: (2024)
BAnG: Bidirectional Anchored Generation for Conditional RNA Design
by: Klypa, Roman, et al.
Published: (2025)
by: Klypa, Roman, et al.
Published: (2025)
Multimodal Alignment and Preference Optimization for Zero-Shot Conditional RNA Generation
by: Klypa, Roman, et al.
Published: (2026)
by: Klypa, Roman, et al.
Published: (2026)
Transformers Provably Learn to Internalize Chain-of-Thought
by: Huang, Yixiao, et al.
Published: (2026)
by: Huang, Yixiao, et al.
Published: (2026)
Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF
by: Zhu, Banghua, et al.
Published: (2024)
by: Zhu, Banghua, et al.
Published: (2024)
dUltra: Ultra-Fast Diffusion Language Models via Reinforcement Learning
by: Chen, Shirui, et al.
Published: (2025)
by: Chen, Shirui, et al.
Published: (2025)
Connections between Schedule-Free Optimizers, AdEMAMix, and Accelerated SGD Variants
by: Morwani, Depen, et al.
Published: (2025)
by: Morwani, Depen, et al.
Published: (2025)
Towards a Theoretical Understanding of the 'Reversal Curse' via Training Dynamics
by: Zhu, Hanlin, et al.
Published: (2024)
by: Zhu, Hanlin, et al.
Published: (2024)
Simulating Environments with Reasoning Models for Agent Training
by: Li, Yuetai, et al.
Published: (2025)
by: Li, Yuetai, et al.
Published: (2025)
MedAgentGym: A Scalable Agentic Training Environment for Code-Centric Reasoning in Biomedical Data Science
by: Xu, Ran, et al.
Published: (2025)
by: Xu, Ran, et al.
Published: (2025)
Physics Steering: Causal Control of Cross-Domain Concepts in a Physics Foundation Model
by: Fear, Rio Alexa, et al.
Published: (2025)
by: Fear, Rio Alexa, et al.
Published: (2025)
Principled Reinforcement Learning with Human Feedback from Pairwise or $K$-wise Comparisons
by: Zhu, Banghua, et al.
Published: (2023)
by: Zhu, Banghua, et al.
Published: (2023)
Reasoning With a Star: A Heliophysics Dataset and Benchmark for Agentic Scientific Reasoning
by: Lee, Kevin, et al.
Published: (2025)
by: Lee, Kevin, et al.
Published: (2025)
CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation
by: Dai, Weinan, et al.
Published: (2026)
by: Dai, Weinan, et al.
Published: (2026)
Learning the Preferences of a Learning Agent
by: Sadek, Karim Abdel, et al.
Published: (2026)
by: Sadek, Karim Abdel, et al.
Published: (2026)
Federated Learning of Socially Appropriate Agent Behaviours in Simulated Home Environments
by: Checker, Saksham, et al.
Published: (2024)
by: Checker, Saksham, et al.
Published: (2024)
Emergence of Superposition: Unveiling the Training Dynamics of Chain of Continuous Thought
by: Zhu, Hanlin, et al.
Published: (2025)
by: Zhu, Hanlin, et al.
Published: (2025)
Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning
by: Wang, Zhaoyang, et al.
Published: (2026)
by: Wang, Zhaoyang, et al.
Published: (2026)
Agentic Proposing: Enhancing Large Language Model Reasoning via Compositional Skill Synthesis
by: Jiao, Zhengbo, et al.
Published: (2026)
by: Jiao, Zhengbo, et al.
Published: (2026)
BAMDP Shaping: a Unified Framework for Intrinsic Motivation and Reward Shaping
by: Lidayan, Aly, et al.
Published: (2024)
by: Lidayan, Aly, et al.
Published: (2024)
ARM: Discovering Agentic Reasoning Modules for Generalizable Multi-Agent Systems
by: Yao, Bohan, et al.
Published: (2025)
by: Yao, Bohan, et al.
Published: (2025)
AgenticPay: A Multi-Agent LLM Negotiation System for Buyer-Seller Transactions
by: Liu, Xianyang, et al.
Published: (2026)
by: Liu, Xianyang, et al.
Published: (2026)
Towards Anytime-Valid Statistical Watermarking
by: Huang, Baihe, et al.
Published: (2026)
by: Huang, Baihe, et al.
Published: (2026)
Generative AI Security: Challenges and Countermeasures
by: Zhu, Banghua, et al.
Published: (2024)
by: Zhu, Banghua, et al.
Published: (2024)
MedAI: Evaluating TxAgent's Therapeutic Agentic Reasoning in the NeurIPS CURE-Bench Competition
by: Cofala, Tim, et al.
Published: (2025)
by: Cofala, Tim, et al.
Published: (2025)
SAGE-32B: Agentic Reasoning via Iterative Distillation
by: Jha, Basab, et al.
Published: (2026)
by: Jha, Basab, et al.
Published: (2026)
Mars: Situated Inductive Reasoning in an Open-World Environment
by: Tang, Xiaojuan, et al.
Published: (2024)
by: Tang, Xiaojuan, et al.
Published: (2024)
Agentic Web: Weaving the Next Web with AI Agents
by: Yang, Yingxuan, et al.
Published: (2025)
by: Yang, Yingxuan, et al.
Published: (2025)
From Bits to Rounds: Parallel Decoding with Exploration for Diffusion Language Models
by: Fu, Hengyu, et al.
Published: (2025)
by: Fu, Hengyu, et al.
Published: (2025)
Imagine-then-Plan: Agent Learning from Adaptive Lookahead with World Models
by: Liu, Youwei, et al.
Published: (2026)
by: Liu, Youwei, et al.
Published: (2026)
Similar Items
-
Understanding the Mechanisms of Fast Hyperparameter Transfer
by: Ghosh, Nikhil, et al.
Published: (2025) -
How Do LLMs Perform Two-Hop Reasoning in Context?
by: Guo, Tianyu, et al.
Published: (2025) -
Generalization or Hallucination? Understanding Out-of-Context Reasoning in Transformers
by: Huang, Yixiao, et al.
Published: (2025) -
Avoiding Catastrophe in Online Learning by Asking for Help
by: Plaut, Benjamin, et al.
Published: (2024) -
GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models
by: Mirzadeh, Iman, et al.
Published: (2024)