The Optimal Token Baseline: Variance Reduction for Long-Horizon LLM-RL
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Yingru, Xu, Jiawei, Li, Ziniu, Liu, Jiacai, Liu, Wei, Tong, Yuxuan, Zheng, Longtao, Xue, Zhenghai, Zhang, Yaxiang, Cai, Tianle, Zhang, Ge, Liu, Qian, Wang, Baoxiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Trust Region Masking for Long-Horizon LLM Reinforcement Learning
von: Li, Yingru, et al.
Veröffentlicht: (2025)
von: Li, Yingru, et al.
Veröffentlicht: (2025)
Dynamic Vocabulary Pruning: Stable LLM-RL by Taming the Tail
von: Li, Yingru, et al.
Veröffentlicht: (2025)
von: Li, Yingru, et al.
Veröffentlicht: (2025)
Beyond Precision: Training-Inference Mismatch is an Optimization Problem and Simple LR Scheduling Fixes It
von: Zhang, Yaxiang, et al.
Veröffentlicht: (2026)
von: Zhang, Yaxiang, et al.
Veröffentlicht: (2026)
A Note on Hybrid Online Reinforcement and Imitation Learning for LLMs: Formulations and Algorithms
von: Li, Yingru, et al.
Veröffentlicht: (2025)
von: Li, Yingru, et al.
Veröffentlicht: (2025)
SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning
von: Xue, Zhenghai, et al.
Veröffentlicht: (2025)
von: Xue, Zhenghai, et al.
Veröffentlicht: (2025)
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents
von: Wang, Jiawei, et al.
Veröffentlicht: (2025)
von: Wang, Jiawei, et al.
Veröffentlicht: (2025)
Dr. Kernel: Reinforcement Learning Done Right for Triton Kernel Generations
von: Liu, Wei, et al.
Veröffentlicht: (2026)
von: Liu, Wei, et al.
Veröffentlicht: (2026)
Scalable Exploration via Ensemble++
von: Li, Yingru, et al.
Veröffentlicht: (2024)
von: Li, Yingru, et al.
Veröffentlicht: (2024)
Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key
von: Wang, Tianle, et al.
Veröffentlicht: (2026)
von: Wang, Tianle, et al.
Veröffentlicht: (2026)
On-Policy RL with Optimal Reward Baseline
von: Hao, Yaru, et al.
Veröffentlicht: (2025)
von: Hao, Yaru, et al.
Veröffentlicht: (2025)
AgentStudio: A Toolkit for Building General Virtual Agents
von: Zheng, Longtao, et al.
Veröffentlicht: (2024)
von: Zheng, Longtao, et al.
Veröffentlicht: (2024)
Sufficient Dimension Reduction via Inverse Conditional Mean or Variance Independence
von: Liu, Jicai, et al.
Veröffentlicht: (2026)
von: Liu, Jicai, et al.
Veröffentlicht: (2026)
Group-in-Group Policy Optimization for LLM Agent Training
von: Feng, Lang, et al.
Veröffentlicht: (2025)
von: Feng, Lang, et al.
Veröffentlicht: (2025)
LongNav-R1: Horizon-Adaptive Multi-Turn RL for Long-Horizon VLA Navigation
von: Hu, Yue, et al.
Veröffentlicht: (2026)
von: Hu, Yue, et al.
Veröffentlicht: (2026)
Horizon Reduction Makes RL Scalable
von: Park, Seohong, et al.
Veröffentlicht: (2025)
von: Park, Seohong, et al.
Veröffentlicht: (2025)
AgentFugue: Agent Scaling for Long-Horizon Tasks through Collective Reasoning
von: Hu, Yuyang, et al.
Veröffentlicht: (2026)
von: Hu, Yuyang, et al.
Veröffentlicht: (2026)
MIND-V: Hierarchical World Model for Long-Horizon Robotic Manipulation with RL-based Physical Alignment
von: Zhang, Ruicheng, et al.
Veröffentlicht: (2025)
von: Zhang, Ruicheng, et al.
Veröffentlicht: (2025)
Elementary Analysis of Policy Gradient Methods
von: Liu, Jiacai, et al.
Veröffentlicht: (2024)
von: Liu, Jiacai, et al.
Veröffentlicht: (2024)
On the Convergence of Policy Mirror Descent with Temporal Difference Evaluation
von: Liu, Jiacai, et al.
Veröffentlicht: (2025)
von: Liu, Jiacai, et al.
Veröffentlicht: (2025)
Logit Dynamics in Softmax Policy Gradient Methods
von: Li, Yingru
Veröffentlicht: (2025)
von: Li, Yingru
Veröffentlicht: (2025)
Probability Tools for Sequential Random Projection
von: Li, Yingru
Veröffentlicht: (2024)
von: Li, Yingru
Veröffentlicht: (2024)
Simple, unified analysis of Johnson-Lindenstrauss with applications
von: Li, Yingru
Veröffentlicht: (2024)
von: Li, Yingru
Veröffentlicht: (2024)
GR-RL: Going Dexterous and Precise for Long-Horizon Robotic Manipulation
von: Li, Yunfei, et al.
Veröffentlicht: (2025)
von: Li, Yunfei, et al.
Veröffentlicht: (2025)
A Scalable Benchmark for Repository-Oriented Long-Horizon Conversational Context Management
von: Liu, Yang, et al.
Veröffentlicht: (2026)
von: Liu, Yang, et al.
Veröffentlicht: (2026)
Divergence-Augmented Policy Optimization
von: Wang, Qing, et al.
Veröffentlicht: (2025)
von: Wang, Qing, et al.
Veröffentlicht: (2025)
Diversity, Variance, and Stability of Root Phenes of Peanut (Arachis hypogaea L.)
von: Lijie Li, et al.
Veröffentlicht: (2024)
von: Lijie Li, et al.
Veröffentlicht: (2024)
Long-context LLMs Struggle with Long In-context Learning
von: Li, Tianle, et al.
Veröffentlicht: (2024)
von: Li, Tianle, et al.
Veröffentlicht: (2024)
SPRINT: Stochastic Performative Prediction With Variance Reduction
von: Xie, Tian, et al.
Veröffentlicht: (2025)
von: Xie, Tian, et al.
Veröffentlicht: (2025)
The Optimal Mean–Variance Selling Problem With Finite Horizon
von: Peter Johnson, et al.
Veröffentlicht: (2026)
von: Peter Johnson, et al.
Veröffentlicht: (2026)
AgentProg: Empowering Long-Horizon GUI Agents with Program-Guided Context Management
von: Tian, Shizuo, et al.
Veröffentlicht: (2025)
von: Tian, Shizuo, et al.
Veröffentlicht: (2025)
The Molecular Structure of Thought: Mapping the Topology of Long Chain-of-Thought Reasoning
von: Chen, Qiguang, et al.
Veröffentlicht: (2026)
von: Chen, Qiguang, et al.
Veröffentlicht: (2026)
Scalable Spatio-Temporal SE(3) Diffusion for Long-Horizon Protein Dynamics
von: Shoghi, Nima, et al.
Veröffentlicht: (2026)
von: Shoghi, Nima, et al.
Veröffentlicht: (2026)
MLorc: Momentum Low-rank Compression for Memory Efficient Large Language Model Adaptation
von: Shen, Wei, et al.
Veröffentlicht: (2025)
von: Shen, Wei, et al.
Veröffentlicht: (2025)
SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks?
von: Lee, Hwiwon, et al.
Veröffentlicht: (2026)
von: Lee, Hwiwon, et al.
Veröffentlicht: (2026)
Hydrogen‐Bond Acceptor and Anion Receptor‐Mediated Regulation of Interfacial Proton Mobility for Long‐Lifespan Aqueous Zinc Batteries
von: Yulong Gao, et al.
Veröffentlicht: (2025)
von: Yulong Gao, et al.
Veröffentlicht: (2025)
Knapsack RL: Unlocking Exploration of LLMs via Optimizing Budget Allocation
von: Li, Ziniu, et al.
Veröffentlicht: (2025)
von: Li, Ziniu, et al.
Veröffentlicht: (2025)
On the Convergence of Projected Policy Gradient for Any Constant Step Sizes
von: Liu, Jiacai, et al.
Veröffentlicht: (2023)
von: Liu, Jiacai, et al.
Veröffentlicht: (2023)
Variance Reduction for the Independent Metropolis Sampler
von: Liu, Siran, et al.
Veröffentlicht: (2024)
von: Liu, Siran, et al.
Veröffentlicht: (2024)
Token-Efficient RL for LLM Reasoning
von: Lee, Alan, et al.
Veröffentlicht: (2025)
von: Lee, Alan, et al.
Veröffentlicht: (2025)
Methyltransferase‐like 14 (METTL14)‐mediated N6‐methyladenosine (m6A) Modification of Forkhead Box Protein 1 (FOXP1) Regulates Trophoblast Inflammation and Function via Transmembrane BAX Inhibitor Motif‐Containing 6 (TMBIM6)
von: Yanhua Liu, et al.
Veröffentlicht: (2025)
von: Yanhua Liu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Trust Region Masking for Long-Horizon LLM Reinforcement Learning
von: Li, Yingru, et al.
Veröffentlicht: (2025) -
Dynamic Vocabulary Pruning: Stable LLM-RL by Taming the Tail
von: Li, Yingru, et al.
Veröffentlicht: (2025) -
Beyond Precision: Training-Inference Mismatch is an Optimization Problem and Simple LR Scheduling Fixes It
von: Zhang, Yaxiang, et al.
Veröffentlicht: (2026) -
A Note on Hybrid Online Reinforcement and Imitation Learning for LLMs: Formulations and Algorithms
von: Li, Yingru, et al.
Veröffentlicht: (2025) -
SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning
von: Xue, Zhenghai, et al.
Veröffentlicht: (2025)