Trust Region Masking for Long-Horizon LLM Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Yingru, Liu, Jiacai, Xu, Jiawei, Tong, Yuxuan, Li, Ziniu, Liu, Qian, Wang, Baoxiang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Optimal Token Baseline: Variance Reduction for Long-Horizon LLM-RL
by: Li, Yingru, et al.
Published: (2026)
by: Li, Yingru, et al.
Published: (2026)
Dynamic Vocabulary Pruning: Stable LLM-RL by Taming the Tail
by: Li, Yingru, et al.
Published: (2025)
by: Li, Yingru, et al.
Published: (2025)
A Note on Hybrid Online Reinforcement and Imitation Learning for LLMs: Formulations and Algorithms
by: Li, Yingru, et al.
Published: (2025)
by: Li, Yingru, et al.
Published: (2025)
Scalable Exploration via Ensemble++
by: Li, Yingru, et al.
Published: (2024)
by: Li, Yingru, et al.
Published: (2024)
Beyond Precision: Training-Inference Mismatch is an Optimization Problem and Simple LR Scheduling Fixes It
by: Zhang, Yaxiang, et al.
Published: (2026)
by: Zhang, Yaxiang, et al.
Published: (2026)
Prior-dependent analysis of posterior sampling reinforcement learning with function approximation
by: Li, Yingru, et al.
Published: (2024)
by: Li, Yingru, et al.
Published: (2024)
Dr. Kernel: Reinforcement Learning Done Right for Triton Kernel Generations
by: Liu, Wei, et al.
Published: (2026)
by: Liu, Wei, et al.
Published: (2026)
Diffusion Model-Based Data Synthesis Aided Federated Semi-Supervised Learning
by: Wang, Zhongwei, et al.
Published: (2025)
by: Wang, Zhongwei, et al.
Published: (2025)
Neural Estimation of Pairwise Mutual Information in Masked Discrete Sequence Models
by: Sharma, Jai, et al.
Published: (2026)
by: Sharma, Jai, et al.
Published: (2026)
Energy Saving for Cell-Free Massive MIMO Networks: A Multi-Agent Deep Reinforcement Learning Approach
by: Wang, Qichen, et al.
Published: (2026)
by: Wang, Qichen, et al.
Published: (2026)
CSGO: Generalized Optimization for Cold Start in Wireless Collaborative Edge LLM Systems
by: Liu, Xuran, et al.
Published: (2025)
by: Liu, Xuran, et al.
Published: (2025)
Information-Theoretic State Variable Selection for Reinforcement Learning
by: Westphal, Charles, et al.
Published: (2024)
by: Westphal, Charles, et al.
Published: (2024)
The Critical Horizon: Inspection Design Principles for Multi-Stage Operations and Deep Reasoning
by: Emadi, Seyed Morteza
Published: (2026)
by: Emadi, Seyed Morteza
Published: (2026)
Environment-Aware Transfer Reinforcement Learning for Sustainable Beam Selection
by: Salami, Dariush, et al.
Published: (2025)
by: Salami, Dariush, et al.
Published: (2025)
Random Aggregate Beamforming for Over-the-Air Federated Learning in Large-Scale Networks
by: Xu, Chunmei, et al.
Published: (2024)
by: Xu, Chunmei, et al.
Published: (2024)
Continual Learning-Aided Super-Resolution Scheme for Channel Reconstruction and Generalization in OFDM Systems
by: Chen, Jianqiao, et al.
Published: (2025)
by: Chen, Jianqiao, et al.
Published: (2025)
Understanding LLM Behaviors via Compression: Data Generation, Knowledge Acquisition and Scaling Laws
by: Pan, Zhixuan, et al.
Published: (2025)
by: Pan, Zhixuan, et al.
Published: (2025)
On the Non-decoupling of Supervised Fine-tuning and Reinforcement Learning in Post-training
by: Niu, Xueyan, et al.
Published: (2026)
by: Niu, Xueyan, et al.
Published: (2026)
ADG: Ambient Diffusion-Guided Dataset Recovery for Corruption-Robust Offline Reinforcement Learning
by: Liu, Zeyuan, et al.
Published: (2025)
by: Liu, Zeyuan, et al.
Published: (2025)
Action-List Reinforcement Learning Syndrome Decoding for Binary Linear Block Codes
by: Taghipour, Milad, et al.
Published: (2025)
by: Taghipour, Milad, et al.
Published: (2025)
Multi-agent Reinforcement Learning for Energy Saving in Multi-Cell Massive MIMO Systems
by: Cai, Tianzhang, et al.
Published: (2024)
by: Cai, Tianzhang, et al.
Published: (2024)
Reinforcement Learning for Long-Horizon Interactive LLM Agents
by: Chen, Kevin, et al.
Published: (2025)
by: Chen, Kevin, et al.
Published: (2025)
Structure-Enhanced Deep Reinforcement Learning for Optimal Transmission Scheduling
by: Chen, Jiazheng, et al.
Published: (2022)
by: Chen, Jiazheng, et al.
Published: (2022)
Dataset Distillation for Offline Reinforcement Learning
by: Light, Jonathan, et al.
Published: (2024)
by: Light, Jonathan, et al.
Published: (2024)
General Information Metrics for Improving AI Model Training Efficiency
by: Xu, Jianfeng, et al.
Published: (2025)
by: Xu, Jianfeng, et al.
Published: (2025)
A Provable Approach for End-to-End Safe Reinforcement Learning
by: Wachi, Akifumi, et al.
Published: (2025)
by: Wachi, Akifumi, et al.
Published: (2025)
Two Birds with One Stone: Multi-Task Semantic Communications Systems over Relay Channel
by: Cao, Yujie, et al.
Published: (2024)
by: Cao, Yujie, et al.
Published: (2024)
Semantic-aware Transmission Scheduling: a Monotonicity-driven Deep Reinforcement Learning Approach
by: Chen, Jiazheng, et al.
Published: (2023)
by: Chen, Jiazheng, et al.
Published: (2023)
Interpretable Diffusion via Information Decomposition
by: Kong, Xianghao, et al.
Published: (2023)
by: Kong, Xianghao, et al.
Published: (2023)
Logit Dynamics in Softmax Policy Gradient Methods
by: Li, Yingru
Published: (2025)
by: Li, Yingru
Published: (2025)
Generative Model-Aided Continual Learning for CSI Feedback in FDD mMIMO-OFDM Systems
by: Liu, Guijun, et al.
Published: (2025)
by: Liu, Guijun, et al.
Published: (2025)
LLMs as Noisy Channels: A Shannon Perspective on Model Capacity and Scaling Laws
by: Ouyang, Xu, et al.
Published: (2026)
by: Ouyang, Xu, et al.
Published: (2026)
MambaJSCC: Adaptive Deep Joint Source-Channel Coding with Generalized State Space Model
by: Wu, Tong, et al.
Published: (2024)
by: Wu, Tong, et al.
Published: (2024)
Optimistic Thompson Sampling for No-Regret Learning in Unknown Games
by: Li, Yingru, et al.
Published: (2024)
by: Li, Yingru, et al.
Published: (2024)
A General Deep Learning Framework for Wireless Resource Allocation under Discrete Constraints
by: Wang, Yikun, et al.
Published: (2026)
by: Wang, Yikun, et al.
Published: (2026)
Provable Privacy Advantages of Decentralized Federated Learning via Distributed Optimization
by: Yu, Wenrui, et al.
Published: (2024)
by: Yu, Wenrui, et al.
Published: (2024)
Semantic Soft Bootstrapping: Long Context Reasoning in LLMs without Reinforcement Learning
by: Mitra, Purbesh, et al.
Published: (2025)
by: Mitra, Purbesh, et al.
Published: (2025)
WDMoE: Wireless Distributed Large Language Models with Mixture of Experts
by: Xue, Nan, et al.
Published: (2024)
by: Xue, Nan, et al.
Published: (2024)
Trajectory-wise Iterative Reinforcement Learning Framework for Auto-bidding
by: Li, Haoming, et al.
Published: (2024)
by: Li, Haoming, et al.
Published: (2024)
RPN 2: On Interdependence Function Learning Towards Unifying and Advancing CNN, RNN, GNN, and Transformer
by: Zhang, Jiawei
Published: (2024)
by: Zhang, Jiawei
Published: (2024)
Similar Items
-
The Optimal Token Baseline: Variance Reduction for Long-Horizon LLM-RL
by: Li, Yingru, et al.
Published: (2026) -
Dynamic Vocabulary Pruning: Stable LLM-RL by Taming the Tail
by: Li, Yingru, et al.
Published: (2025) -
A Note on Hybrid Online Reinforcement and Imitation Learning for LLMs: Formulations and Algorithms
by: Li, Yingru, et al.
Published: (2025) -
Scalable Exploration via Ensemble++
by: Li, Yingru, et al.
Published: (2024) -
Beyond Precision: Training-Inference Mismatch is an Optimization Problem and Simple LR Scheduling Fixes It
by: Zhang, Yaxiang, et al.
Published: (2026)