EntroPIC: Towards Stable Long-Term Training of LLMs via Entropy Stabilization with Proportional-Integral Control
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Kai, Xu, Xin, Chen, Yangkun, Liu, Weijie, Lyu, Jiafei, Lin, Zichuan, Ye, Deheng, Yang, Saiyong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Debiased Model-based Representations for Sample-efficient Continuous Control
by: Lyu, Jiafei, et al.
Published: (2026)
by: Lyu, Jiafei, et al.
Published: (2026)
Novelty-Guided Data Reuse for Efficient and Diversified Multi-Agent Reinforcement Learning
by: Chen, Yangkun, et al.
Published: (2024)
by: Chen, Yangkun, et al.
Published: (2024)
Cross-Domain Offline Policy Adaptation via Selective Transition Correction
by: Yan, Mengbei, et al.
Published: (2026)
by: Yan, Mengbei, et al.
Published: (2026)
Composition-RL: Compose Your Verifiable Prompts for Reinforcement Learning of Large Language Models
by: Xu, Xin, et al.
Published: (2026)
by: Xu, Xin, et al.
Published: (2026)
Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models
by: Qu, Yun, et al.
Published: (2026)
by: Qu, Yun, et al.
Published: (2026)
AdaptVision: Efficient Vision-Language Models via Adaptive Visual Acquisition
by: Lin, Zichuan, et al.
Published: (2025)
by: Lin, Zichuan, et al.
Published: (2025)
SeeNav-Agent: Enhancing Vision-Language Navigation with Visual Prompt and Step-Level Policy Optimization
by: Wang, Zhengcheng, et al.
Published: (2025)
by: Wang, Zhengcheng, et al.
Published: (2025)
HiRO-Nav: Hybrid ReasOning Enables Efficient Embodied Navigation
by: Zhao, He, et al.
Published: (2026)
by: Zhao, He, et al.
Published: (2026)
EntroCoT: Enhancing Chain-of-Thought via Adaptive Entropy-Guided Segmentation
by: Li, Zihang, et al.
Published: (2026)
by: Li, Zihang, et al.
Published: (2026)
The Lifecycle Principle: Stabilizing Dynamic Neural Networks with State Memory
by: Yang, Zichuan
Published: (2025)
by: Yang, Zichuan
Published: (2025)
Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
by: Yang, Wenkai, et al.
Published: (2026)
by: Yang, Wenkai, et al.
Published: (2026)
ProAct: Agentic Lookahead in Interactive Environments
by: Yu, Yangbin, et al.
Published: (2026)
by: Yu, Yangbin, et al.
Published: (2026)
UI-Voyager: A Self-Evolving GUI Agent Learning via Failed Experience
by: Lin, Zichuan, et al.
Published: (2026)
by: Lin, Zichuan, et al.
Published: (2026)
Temporal Difference Learning with Constrained Initial Representations
by: Lyu, Jiafei, et al.
Published: (2026)
by: Lyu, Jiafei, et al.
Published: (2026)
Stability Analysis of Proportional-Integral AQM Controllers Supporting TCP Flows
by: Daniel Melchor Aguilar
Published: (2007)
by: Daniel Melchor Aguilar
Published: (2007)
Boosting Vulnerability Detection of LLMs via Curriculum Preference Optimization with Synthetic Reasoning Data
by: Wen, Xin-Cheng, et al.
Published: (2025)
by: Wen, Xin-Cheng, et al.
Published: (2025)
PASK: Toward Intent-Aware Proactive Agents with Long-Term Memory
by: Xie, Zhifei, et al.
Published: (2026)
by: Xie, Zhifei, et al.
Published: (2026)
Learning Versatile Skills with Curriculum Masking
by: Tang, Yao, et al.
Published: (2024)
by: Tang, Yao, et al.
Published: (2024)
PROF: An LLM-based Reward Code Preference Optimization Framework for Offline Imitation Learning
by: Sun, Shengjie, et al.
Published: (2025)
by: Sun, Shengjie, et al.
Published: (2025)
EntroAD: Structural Entropy-Guided Prompt Adaptation for Zero-Shot Anomaly Detection
by: Zhao, Xinyu, et al.
Published: (2026)
by: Zhao, Xinyu, et al.
Published: (2026)
ORBIT: On-policy Exploration-Exploitation for Controllable Multi-Budget Reasoning
by: Liang, Kun, et al.
Published: (2026)
by: Liang, Kun, et al.
Published: (2026)
A Novel Proportional‐Integral Synergetic Controller for the Onboard Microgrids With Stability Analysis
by: Fangyuan Li, et al.
Published: (2025)
by: Fangyuan Li, et al.
Published: (2025)
Exploration and Anti-Exploration with Distributional Random Network Distillation
by: Yang, Kai, et al.
Published: (2024)
by: Yang, Kai, et al.
Published: (2024)
A Two-stage Reinforcement Learning-based Approach for Multi-entity Task Allocation
by: Gong, Aicheng, et al.
Published: (2024)
by: Gong, Aicheng, et al.
Published: (2024)
Multi-Tier UAV Edge Computing for Low Altitude Networks Towards Long-Term Energy Stability
by: Ye, Yufei, et al.
Published: (2025)
by: Ye, Yufei, et al.
Published: (2025)
Multi-Tier UAV Edge Computing Towards Long-Term Energy Stability for Low Altitude Networks
by: Ye, Yufei, et al.
Published: (2026)
by: Ye, Yufei, et al.
Published: (2026)
Analytical Design of Proportional Integral AQM Controllers for Enhanced Stability and Performance in TCP Systems
by: Ouassim Menacer, et al.
Published: (2025)
by: Ouassim Menacer, et al.
Published: (2025)
EntroLLM: Entropy Encoded Weight Compression for Efficient Large Language Model Inference on Edge Devices
by: Sanyal, Arnab, et al.
Published: (2025)
by: Sanyal, Arnab, et al.
Published: (2025)
EntroLnn: Entropy-Guided Liquid Neural Networks for Operando Refinement of Battery Capacity Fade Trajectories
by: Li, Wei, et al.
Published: (2026)
by: Li, Wei, et al.
Published: (2026)
Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex
by: Qu, Yun, et al.
Published: (2026)
by: Qu, Yun, et al.
Published: (2026)
Thinking-Free Policy Initialization Makes Distilled Reasoning Models More Effective and Efficient Reasoners
by: Xu, Xin, et al.
Published: (2025)
by: Xu, Xin, et al.
Published: (2025)
MetaColloc: Optimization-Free PDE Solving via Meta-Learned Basis Functions
by: Yang, Zichuan
Published: (2026)
by: Yang, Zichuan
Published: (2026)
CausalMACE: Causality Empowered Multi-Agents in Minecraft Cooperative Tasks
by: Chai, Qi, et al.
Published: (2025)
by: Chai, Qi, et al.
Published: (2025)
ADWIN: Adaptive Windows for Horizon-Aware On-Policy Distillation
by: Liang, Kun, et al.
Published: (2026)
by: Liang, Kun, et al.
Published: (2026)
LaSeR: Reinforcement Learning with Last-Token Self-Rewarding
by: Yang, Wenkai, et al.
Published: (2025)
by: Yang, Wenkai, et al.
Published: (2025)
EntroCut: Entropy-Guided Adaptive Truncation for Efficient Chain-of-Thought Reasoning in Small-scale Large Reasoning Models
by: Yan, Hongxi, et al.
Published: (2026)
by: Yan, Hongxi, et al.
Published: (2026)
Mechanically Compliant and Stable Hydrogel Interfaces for Long‐Term Dynamic Pressure Detection Toward Intelligent Sports Training
by: Hongnan Zhu, et al.
Published: (2025)
by: Hongnan Zhu, et al.
Published: (2025)
Homogeneous Proportional-Integral-Derivative Controller in Mobile Robotic Manipulators
by: Luna, Luis, et al.
Published: (2025)
by: Luna, Luis, et al.
Published: (2025)
Do Not Step Into the Same River Twice: Learning to Reason from Trial and Error
by: Tang, Chenming, et al.
Published: (2025)
by: Tang, Chenming, et al.
Published: (2025)
Democratizing Tool Learning with Environments Fully Simulated by a Free 8B Language Model
by: Tang, Chenming, et al.
Published: (2026)
by: Tang, Chenming, et al.
Published: (2026)
Similar Items
-
Debiased Model-based Representations for Sample-efficient Continuous Control
by: Lyu, Jiafei, et al.
Published: (2026) -
Novelty-Guided Data Reuse for Efficient and Diversified Multi-Agent Reinforcement Learning
by: Chen, Yangkun, et al.
Published: (2024) -
Cross-Domain Offline Policy Adaptation via Selective Transition Correction
by: Yan, Mengbei, et al.
Published: (2026) -
Composition-RL: Compose Your Verifiable Prompts for Reinforcement Learning of Large Language Models
by: Xu, Xin, et al.
Published: (2026) -
Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models
by: Qu, Yun, et al.
Published: (2026)