Scaling Behaviors of LLM Reinforcement Learning Post-Training: An Empirical Study in Mathematical Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Tan, Zelin, Geng, Hejia, Yu, Xiaohang, Zhang, Mulei, Wan, Guancheng, Zhou, Yifan, He, Qiang, Xue, Xiangyuan, Zhou, Heng, Fan, Yutao, Li, Zhongzhi, Zhang, Zaibin, Zhang, Guibin, Zhang, Chen, Yin, Zhenfei, Torr, Philip, Bai, Lei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
by: Zhang, Guibin, et al.
Published: (2025)
by: Zhang, Guibin, et al.
Published: (2025)
CoMAS: Co-Evolving Multi-Agent Systems via Interaction Rewards
by: Xue, Xiangyuan, et al.
Published: (2025)
by: Xue, Xiangyuan, et al.
Published: (2025)
PAPO: Stabilizing Rubric Integration Training via Decoupled Advantage Normalization
by: Tan, Zelin, et al.
Published: (2026)
by: Tan, Zelin, et al.
Published: (2026)
ReSo: A Reward-driven Self-organizing LLM-based Multi-Agent System for Reasoning Tasks
by: Zhou, Heng, et al.
Published: (2025)
by: Zhou, Heng, et al.
Published: (2025)
StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction
by: Xue, Xiangyuan, et al.
Published: (2026)
by: Xue, Xiangyuan, et al.
Published: (2026)
LatentEvolve: Self-Evolving Test-Time Scaling in Latent Space
by: Zhang, Guibin, et al.
Published: (2025)
by: Zhang, Guibin, et al.
Published: (2025)
Diagnose, Localize, Align: A Full-Stack Framework for Reliable LLM Multi-Agent Systems under Instruction Conflicts
by: Wan, Guancheng, et al.
Published: (2025)
by: Wan, Guancheng, et al.
Published: (2025)
EvoFlow: Evolving Diverse Agentic Workflows On The Fly
by: Zhang, Guibin, et al.
Published: (2025)
by: Zhang, Guibin, et al.
Published: (2025)
LiveSearchBench: An Automatically Constructed Benchmark for Retrieval and Reasoning over Dynamic Knowledge
by: Zhou, Heng, et al.
Published: (2025)
by: Zhou, Heng, et al.
Published: (2025)
SUDP: Secret-Use Delegation Protocol for Agentic Systems
by: Yu, Xiaohang, et al.
Published: (2026)
by: Yu, Xiaohang, et al.
Published: (2026)
Behavioral Consistency Validation for LLM Agents: An Analysis of Trading-Style Switching through Stock-Market Simulation
by: Li, Zeping, et al.
Published: (2026)
by: Li, Zeping, et al.
Published: (2026)
CoEnv: Driving Embodied Multi-Agent Collaboration via Compositional Environment
by: Kang, Li, et al.
Published: (2026)
by: Kang, Li, et al.
Published: (2026)
Composing Recurrent Spiking Neural Networks using Locally-Recurrent Motifs and Risk-Mitigating Architectural Optimization
by: Zhang, Wenrui, et al.
Published: (2021)
by: Zhang, Wenrui, et al.
Published: (2021)
G-Memory: Tracing Hierarchical Memory for Multi-Agent Systems
by: Zhang, Guibin, et al.
Published: (2025)
by: Zhang, Guibin, et al.
Published: (2025)
VIKI-R: Coordinating Embodied Multi-Agent Cooperation via Reinforcement Learning
by: Kang, Li, et al.
Published: (2025)
by: Kang, Li, et al.
Published: (2025)
Multi-Task Offloading via Graph Neural Networks in Heterogeneous Multi-access Edge Computing
by: Ma, Mulei
Published: (2023)
by: Ma, Mulei
Published: (2023)
BEV-IO: Enhancing Bird's-Eye-View 3D Detection with Instance Occupancy
by: Zhang, Zaibin, et al.
Published: (2023)
by: Zhang, Zaibin, et al.
Published: (2023)
Ego to World: Collaborative Spatial Reasoning in Embodied Systems via Reinforcement Learning
by: Zhou, Heng, et al.
Published: (2026)
by: Zhou, Heng, et al.
Published: (2026)
Reading $\neq$ Seeing: Diagnosing and Closing the Typography Gap in Vision-Language Models
by: Zhou, Heng, et al.
Published: (2026)
by: Zhou, Heng, et al.
Published: (2026)
Deconstructing the Burden of History: Gender Internalisation and the Epistemology of Illegitimate Tasks in Nursing Practice
by: Zilin Zhao, et al.
Published: (2025)
by: Zilin Zhao, et al.
Published: (2025)
TodoEvolve: Learning to Architect Agent Planning Systems
by: Liu, Jiaxi, et al.
Published: (2026)
by: Liu, Jiaxi, et al.
Published: (2026)
Fast Computation for the Forest Matrix of an Evolving Graph
by: Sun, Haoxin, et al.
Published: (2024)
by: Sun, Haoxin, et al.
Published: (2024)
Privacy-Enhancing Paradigms within Federated Multi-Agent Systems
by: Shi, Zitong, et al.
Published: (2025)
by: Shi, Zitong, et al.
Published: (2025)
MasRouter: Learning to Route LLMs for Multi-Agent Systems
by: Yue, Yanwei, et al.
Published: (2025)
by: Yue, Yanwei, et al.
Published: (2025)
EPIC: A Lightweight LiDAR-Based UAV Exploration Framework for Large-Scale Scenarios
by: Geng, Shuang, et al.
Published: (2024)
by: Geng, Shuang, et al.
Published: (2024)
HoSNN: Adversarially-Robust Homeostatic Spiking Neural Networks with Adaptive Firing Thresholds
by: Geng, Hejia, et al.
Published: (2023)
by: Geng, Hejia, et al.
Published: (2023)
A-MapReduce: Executing Wide Search via Agentic MapReduce
by: Chen, Mingju, et al.
Published: (2026)
by: Chen, Mingju, et al.
Published: (2026)
HarnessForge: Joint Harness and Policy Evolution for Adaptive Agent Systems
by: Chen, Mingju, et al.
Published: (2026)
by: Chen, Mingju, et al.
Published: (2026)
Coupling Analytical Investigation on the Response of an Existing Tunnel Subjected to Underpassing Shield Tunneling
by: Kai Zhang, et al.
Published: (2025)
by: Kai Zhang, et al.
Published: (2025)
Efficient Algorithms for Minimizing the Kirchhoff Index via Adding Edges
by: Zhou, Xiaotian, et al.
Published: (2025)
by: Zhou, Xiaotian, et al.
Published: (2025)
Think3D: Thinking with Space for Spatial Reasoning
by: Zhang, Zaibin, et al.
Published: (2026)
by: Zhang, Zaibin, et al.
Published: (2026)
Eigen-1: Adaptive Multi-Agent Refinement with Monitor-Based RAG for Scientific Reasoning
by: Tang, Xiangru, et al.
Published: (2025)
by: Tang, Xiangru, et al.
Published: (2025)
Reducibility of ultra‐differentiable quasi‐periodic linear systems
by: Xiangyuan Zhang, et al.
Published: (2025)
by: Xiangyuan Zhang, et al.
Published: (2025)
Behavior and Sublinear Algorithm for Opinion Disagreement on Noisy Social Networks
by: Xu, Wanyue, et al.
Published: (2026)
by: Xu, Wanyue, et al.
Published: (2026)
G-Safeguard: A Topology-Guided Security Lens and Treatment on LLM-based Multi-agent Systems
by: Wang, Shilong, et al.
Published: (2025)
by: Wang, Shilong, et al.
Published: (2025)
Select-then-Solve: Paradigm Routing as Inference-Time Optimization for LLM Agents
by: Zhou, Heng, et al.
Published: (2026)
by: Zhou, Heng, et al.
Published: (2026)
Fast Maximization of Current Flow Group Closeness Centrality
by: Xia, Haisong, et al.
Published: (2025)
by: Xia, Haisong, et al.
Published: (2025)
ProbeWalk: Fast Estimation of Biharmonic Distance on Graphs via Probe-Driven Random Walks
by: Zheng, Dehong, et al.
Published: (2025)
by: Zheng, Dehong, et al.
Published: (2025)
Fast Computation of Kemeny's Constant for Directed Graphs
by: Xia, Haisong, et al.
Published: (2024)
by: Xia, Haisong, et al.
Published: (2024)
AD-H: Language-guided Autonomous Driving with Hierarchical Agents
by: Zhang, Zaibin, et al.
Published: (2024)
by: Zhang, Zaibin, et al.
Published: (2024)
Similar Items
-
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
by: Zhang, Guibin, et al.
Published: (2025) -
CoMAS: Co-Evolving Multi-Agent Systems via Interaction Rewards
by: Xue, Xiangyuan, et al.
Published: (2025) -
PAPO: Stabilizing Rubric Integration Training via Decoupled Advantage Normalization
by: Tan, Zelin, et al.
Published: (2026) -
ReSo: A Reward-driven Self-organizing LLM-based Multi-Agent System for Reasoning Tasks
by: Zhou, Heng, et al.
Published: (2025) -
StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction
by: Xue, Xiangyuan, et al.
Published: (2026)