Temperature as a Meta-Policy: Adaptive Temperature in LLM Reinforcement Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dang, Haoran, Lan, Cuiling, Wan, Hai, Zhao, Xibin, Lu, Yan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Adversarial Contrastive Learning for LLM Quantization Attacks
von: Song, Dinghong, et al.
Veröffentlicht: (2026)
von: Song, Dinghong, et al.
Veröffentlicht: (2026)
Internalize the Temperature: On-Policy Self-Distillation as Policy Reheater for Reinforcement Learning
von: Yang, Xuewei, et al.
Veröffentlicht: (2026)
von: Yang, Xuewei, et al.
Veröffentlicht: (2026)
Periodic Asynchrony: An On-Policy Approach for Accelerating LLM Reinforcement Learning
von: Lu, Jian
Veröffentlicht: (2025)
von: Lu, Jian
Veröffentlicht: (2025)
Adaptive Advantage-Guided Policy Regularization for Offline Reinforcement Learning
von: Liu, Tenglong, et al.
Veröffentlicht: (2024)
von: Liu, Tenglong, et al.
Veröffentlicht: (2024)
EPO: Entropy-regularized Policy Optimization for LLM Agents Reinforcement Learning
von: Xu, Wujiang, et al.
Veröffentlicht: (2025)
von: Xu, Wujiang, et al.
Veröffentlicht: (2025)
AdaMemento: Adaptive Memory-Assisted Policy Optimization for Reinforcement Learning
von: Yan, Renye, et al.
Veröffentlicht: (2024)
von: Yan, Renye, et al.
Veröffentlicht: (2024)
Adaptive Scaling of Policy Constraints for Offline Reinforcement Learning
von: Jing, Tan, et al.
Veröffentlicht: (2025)
von: Jing, Tan, et al.
Veröffentlicht: (2025)
Playing Non-Embedded Card-Based Games with Reinforcement Learning
von: Wu, Tianyang, et al.
Veröffentlicht: (2025)
von: Wu, Tianyang, et al.
Veröffentlicht: (2025)
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search
von: Hou, Zhenyu, et al.
Veröffentlicht: (2025)
von: Hou, Zhenyu, et al.
Veröffentlicht: (2025)
Towards an Information Theoretic Framework of Context-Based Offline Meta-Reinforcement Learning
von: Li, Lanqing, et al.
Veröffentlicht: (2024)
von: Li, Lanqing, et al.
Veröffentlicht: (2024)
Revisiting Graph-Based Fraud Detection in Sight of Heterophily and Spectrum
von: Xu, Fan, et al.
Veröffentlicht: (2023)
von: Xu, Fan, et al.
Veröffentlicht: (2023)
LLM-Guided Reinforcement Learning: Addressing Training Bottlenecks through Policy Modulation
von: Tan, Heng, et al.
Veröffentlicht: (2025)
von: Tan, Heng, et al.
Veröffentlicht: (2025)
Adaptive Policy Synchronization for Scalable Reinforcement Learning
von: Lafuente-Mercado, Rodney
Veröffentlicht: (2025)
von: Lafuente-Mercado, Rodney
Veröffentlicht: (2025)
Learning When to Switch: Adaptive Policy Selection via Reinforcement Learning
von: Tava, Chris
Veröffentlicht: (2025)
von: Tava, Chris
Veröffentlicht: (2025)
LESS: Efficient Log Storage System Based on Learned Model and Minimum Attribute Tree
von: Cheng, Zhiyang, et al.
Veröffentlicht: (2024)
von: Cheng, Zhiyang, et al.
Veröffentlicht: (2024)
GLADformer: A Mixed Perspective for Graph-level Anomaly Detection
von: Xu, Fan, et al.
Veröffentlicht: (2024)
von: Xu, Fan, et al.
Veröffentlicht: (2024)
Meta-Learning and Meta-Reinforcement Learning -- Tracing the Path towards DeepMind's Adaptive Agent
von: Hoppmann, Björn, et al.
Veröffentlicht: (2026)
von: Hoppmann, Björn, et al.
Veröffentlicht: (2026)
Actor-Accelerated Policy Dual Averaging for Reinforcement Learning in Continuous Action Spaces
von: Gao, Ji, et al.
Veröffentlicht: (2026)
von: Gao, Ji, et al.
Veröffentlicht: (2026)
Reinforcement Learning of Adaptive Acquisition Policies for Inverse Problems
von: Silvestri, Gianluigi, et al.
Veröffentlicht: (2024)
von: Silvestri, Gianluigi, et al.
Veröffentlicht: (2024)
Adaptive Multi-Agent Deep Reinforcement Learning for Timely Healthcare Interventions
von: Shaik, Thanveer, et al.
Veröffentlicht: (2023)
von: Shaik, Thanveer, et al.
Veröffentlicht: (2023)
Using Temperature Sampling to Effectively Train Robot Learning Policies on Imbalanced Datasets
von: Patil, Basavasagar, et al.
Veröffentlicht: (2025)
von: Patil, Basavasagar, et al.
Veröffentlicht: (2025)
Improving Domain Generalization in Contrastive Learning using Adaptive Temperature Control
von: Lewis, Robert, et al.
Veröffentlicht: (2026)
von: Lewis, Robert, et al.
Veröffentlicht: (2026)
On the Reuse Bias in Off-Policy Reinforcement Learning
von: Ying, Chengyang, et al.
Veröffentlicht: (2022)
von: Ying, Chengyang, et al.
Veröffentlicht: (2022)
Scalable On-Policy Reinforcement Learning via Adaptive Batch Scaling
von: Park, Jongchan
Veröffentlicht: (2026)
von: Park, Jongchan
Veröffentlicht: (2026)
Offline Policy Evaluation for Reinforcement Learning with Adaptively Collected Data
von: Madhow, Sunil, et al.
Veröffentlicht: (2023)
von: Madhow, Sunil, et al.
Veröffentlicht: (2023)
Constraint-Adaptive Policy Switching for Offline Safe Reinforcement Learning
von: Chemingui, Yassine, et al.
Veröffentlicht: (2024)
von: Chemingui, Yassine, et al.
Veröffentlicht: (2024)
Look Inward to Explore Outward: Learning Temperature Policy from LLM Internal States via Hierarchical RL
von: Zhou, Yixiao, et al.
Veröffentlicht: (2026)
von: Zhou, Yixiao, et al.
Veröffentlicht: (2026)
Adaptive Simulation Experiment for LLM Policy Optimization
von: Hu, Mingjie, et al.
Veröffentlicht: (2026)
von: Hu, Mingjie, et al.
Veröffentlicht: (2026)
Learning to Optimize for Reinforcement Learning
von: Lan, Qingfeng, et al.
Veröffentlicht: (2023)
von: Lan, Qingfeng, et al.
Veröffentlicht: (2023)
A Phone-based Distributed Ambient Temperature Measurement System with An Efficient Label-free Automated Training Strategy
von: Chen, Dayin, et al.
Veröffentlicht: (2024)
von: Chen, Dayin, et al.
Veröffentlicht: (2024)
Boosting Continuous Control with Consistency Policy
von: Chen, Yuhui, et al.
Veröffentlicht: (2023)
von: Chen, Yuhui, et al.
Veröffentlicht: (2023)
Evolving Diffusion and Flow Matching Policies for Online Reinforcement Learning
von: Zhang, Chubin, et al.
Veröffentlicht: (2025)
von: Zhang, Chubin, et al.
Veröffentlicht: (2025)
Offline Meta-Reinforcement Learning with Flow-Based Task Inference and Adaptive Correction of Feature Overgeneralization
von: Wang, Min, et al.
Veröffentlicht: (2026)
von: Wang, Min, et al.
Veröffentlicht: (2026)
Adaptive Policy Learning to Additional Tasks
von: Hao, Wenjian, et al.
Veröffentlicht: (2023)
von: Hao, Wenjian, et al.
Veröffentlicht: (2023)
Policy Zooming: Adaptive Discretization-based Infinite-Horizon Average-Reward Reinforcement Learning
von: Kar, Avik, et al.
Veröffentlicht: (2024)
von: Kar, Avik, et al.
Veröffentlicht: (2024)
Graph-enabled Reinforcement Learning for Time Series Forecasting with Adaptive Intelligence
von: Shaik, Thanveer, et al.
Veröffentlicht: (2023)
von: Shaik, Thanveer, et al.
Veröffentlicht: (2023)
A Self-Attentive Meta-Optimizer with Group-Adaptive Learning Rates and Weight Decay
von: Zhao, JiangBo, et al.
Veröffentlicht: (2026)
von: Zhao, JiangBo, et al.
Veröffentlicht: (2026)
FM-IRL: Flow-Matching for Reward Modeling and Policy Regularization in Reinforcement Learning
von: Wan, Zhenglin, et al.
Veröffentlicht: (2025)
von: Wan, Zhenglin, et al.
Veröffentlicht: (2025)
GHPO: Adaptive Guidance for Stable and Efficient LLM Reinforcement Learning
von: Liu, Ziru, et al.
Veröffentlicht: (2025)
von: Liu, Ziru, et al.
Veröffentlicht: (2025)
Private Optimal Inventory Policy Learning for Feature-based Newsvendor with Unknown Demand
von: Zhao, Tuoyi, et al.
Veröffentlicht: (2024)
von: Zhao, Tuoyi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Adversarial Contrastive Learning for LLM Quantization Attacks
von: Song, Dinghong, et al.
Veröffentlicht: (2026) -
Internalize the Temperature: On-Policy Self-Distillation as Policy Reheater for Reinforcement Learning
von: Yang, Xuewei, et al.
Veröffentlicht: (2026) -
Periodic Asynchrony: An On-Policy Approach for Accelerating LLM Reinforcement Learning
von: Lu, Jian
Veröffentlicht: (2025) -
Adaptive Advantage-Guided Policy Regularization for Offline Reinforcement Learning
von: Liu, Tenglong, et al.
Veröffentlicht: (2024) -
EPO: Entropy-regularized Policy Optimization for LLM Agents Reinforcement Learning
von: Xu, Wujiang, et al.
Veröffentlicht: (2025)