VEPO: Variable Entropy Policy Optimization for Low-Resource Language Foundation Models
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Chonghan, Du, Yimin, An, Qi, He, Xin, Zhai, Cunqi, Tan, Fei, Lin, Weijia, Gong, Xiaochun, Deng, Yongchao, Jia, Shousheng, Zhang, Xiangzheng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
360-LLaMA-Factory: Plug & Play Sequence Parallelism for Long Post-Training
by: Zou, Haosheng, et al.
Published: (2025)
by: Zou, Haosheng, et al.
Published: (2025)
Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
by: Wen, Liang, et al.
Published: (2025)
by: Wen, Liang, et al.
Published: (2025)
Light-IF: Endowing LLMs with Generalizable Reasoning via Preview and Self-Checking for Complex Instruction Following
by: Wang, Chenyang, et al.
Published: (2025)
by: Wang, Chenyang, et al.
Published: (2025)
The role of lipoylation in mitochondrial adaptation to methionine restriction
by: Jingyuan Xue, et al.
Published: (2024)
by: Jingyuan Xue, et al.
Published: (2024)
Phospholipid Biosynthesis: An Unforeseen Modulator of Nuclear Metabolism
by: Hong Qiu, et al.
Published: (2025)
by: Hong Qiu, et al.
Published: (2025)
Machine Learning Enhanced Multi-Factor Quantitative Trading: A Cross-Sectional Portfolio Optimization Approach with Bias Correction
by: Du, Yimin
Published: (2025)
by: Du, Yimin
Published: (2025)
Foundation Models for Low-Resource Language Education (Vision Paper)
by: Ding, Zhaojun, et al.
Published: (2024)
by: Ding, Zhaojun, et al.
Published: (2024)
Teola: Towards End-to-End Optimization of LLM-based Applications
by: Tan, Xin, et al.
Published: (2024)
by: Tan, Xin, et al.
Published: (2024)
Water‐Energy‐Carbon Nexus Within the Urban Eco‐Transformation of the Beijing‐Tianjin‐Hebei Region
by: Yifei Wang, et al.
Published: (2025)
by: Yifei Wang, et al.
Published: (2025)
SRAP-Agent: Simulating and Optimizing Scarce Resource Allocation Policy with LLM-based Agent
by: Ji, Jiarui, et al.
Published: (2024)
by: Ji, Jiarui, et al.
Published: (2024)
Augmenting Document-level Relation Extraction with Efficient Multi-Supervision
by: Lin, Xiangyu, et al.
Published: (2024)
by: Lin, Xiangyu, et al.
Published: (2024)
DCPO: Dynamic Clipping Policy Optimization
by: Yang, Shihui, et al.
Published: (2025)
by: Yang, Shihui, et al.
Published: (2025)
Intuitionistic Fuzzy Sets for Large Language Model Data Annotation: A Novel Approach to Side-by-Side Preference Labeling
by: Du, Yimin
Published: (2025)
by: Du, Yimin
Published: (2025)
Deep Learning Enhanced Multi-Day Turnover Quantitative Trading Algorithm for Chinese A-Share Market
by: Du, Yimin
Published: (2025)
by: Du, Yimin
Published: (2025)
Memory-Efficient FastText: A Comprehensive Approach Using Double-Array Trie Structures and Mark-Compact Memory Management
by: Du, Yimin
Published: (2025)
by: Du, Yimin
Published: (2025)
ARC: Active and Reflection-driven Context Management for Long-Horizon Information Seeking Agents
by: Yao, Yilun, et al.
Published: (2026)
by: Yao, Yilun, et al.
Published: (2026)
Uncertainty Under the Curve: A Sequence-Level Entropy Area Metric for Reasoning LLM
by: Zhu, Yongfu, et al.
Published: (2025)
by: Zhu, Yongfu, et al.
Published: (2025)
Entropy-Regularized Token-Level Policy Optimization for Language Agent Reinforcement
by: Wen, Muning, et al.
Published: (2024)
by: Wen, Muning, et al.
Published: (2024)
Beyond Static Alignment: Hierarchical Policy Control for LLM Safety via Risk-Aware Chain-of-Thought
by: Si, Jianfeng, et al.
Published: (2026)
by: Si, Jianfeng, et al.
Published: (2026)
Agentic Entropy-Balanced Policy Optimization
by: Dong, Guanting, et al.
Published: (2025)
by: Dong, Guanting, et al.
Published: (2025)
Relative Entropy Pathwise Policy Optimization
by: Voelcker, Claas, et al.
Published: (2025)
by: Voelcker, Claas, et al.
Published: (2025)
Efficient Switchable Safety Control in LLMs via Magic-Token-Guided Co-Training
by: Si, Jianfeng, et al.
Published: (2025)
by: Si, Jianfeng, et al.
Published: (2025)
Joint Optimization of Prompt Security and System Performance in Edge-Cloud LLM Systems
by: Huang, Haiyang, et al.
Published: (2025)
by: Huang, Haiyang, et al.
Published: (2025)
Low-Resource Vision Challenges for Foundation Models
by: Zhang, Yunhua, et al.
Published: (2024)
by: Zhang, Yunhua, et al.
Published: (2024)
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization
by: Xu, Huimin, et al.
Published: (2026)
by: Xu, Huimin, et al.
Published: (2026)
Cognitive Edge Computing: A Comprehensive Survey on Optimizing Large Models and AI Agents for Pervasive Deployment
by: Wang, Xubin, et al.
Published: (2025)
by: Wang, Xubin, et al.
Published: (2025)
One-Step Flow Policy: Self-Distillation for Fast Visuomotor Policies
by: Li, Shaolong, et al.
Published: (2026)
by: Li, Shaolong, et al.
Published: (2026)
Low-rank Orthogonalization for Large-scale Matrix Optimization with Applications to Foundation Model Training
by: He, Chuan, et al.
Published: (2025)
by: He, Chuan, et al.
Published: (2025)
When Maximum Entropy Misleads Policy Optimization
by: Zhang, Ruipeng, et al.
Published: (2025)
by: Zhang, Ruipeng, et al.
Published: (2025)
ESPO: Entropy Importance Sampling Policy Optimization
by: Sheng, Yuepeng, et al.
Published: (2025)
by: Sheng, Yuepeng, et al.
Published: (2025)
VSPO: Vector-Steered Policy Optimization for Behavioral Control
by: Zhang, Xuechen, et al.
Published: (2026)
by: Zhang, Xuechen, et al.
Published: (2026)
Animating the Past: Reconstruct Trilobite via Video Generation
by: Wu, Xiaoran, et al.
Published: (2024)
by: Wu, Xiaoran, et al.
Published: (2024)
Diffusion Posterior Sampler for Hyperspectral Unmixing with Spectral Variability Modeling
by: Zhu, Yimin, et al.
Published: (2025)
by: Zhu, Yimin, et al.
Published: (2025)
OmniArch: Building Foundation Model For Scientific Computing
by: Chen, Tianyu, et al.
Published: (2024)
by: Chen, Tianyu, et al.
Published: (2024)
Optimistic Model Rollouts for Pessimistic Offline Policy Optimization
by: Zhai, Yuanzhao, et al.
Published: (2024)
by: Zhai, Yuanzhao, et al.
Published: (2024)
Frontier: Towards Comprehensive and Accurate LLM Inference Simulation
by: Feng, Yicheng, et al.
Published: (2026)
by: Feng, Yicheng, et al.
Published: (2026)
Architecting Gradient Hierarchically Porous Catalyst via Negative Mixing Enthalpy High‐Entropy Alloy for Durable Water Splitting at Ampere‐Level Current Density
by: Qiqin Zhang, et al.
Published: (2024)
by: Qiqin Zhang, et al.
Published: (2024)
Lifting Data-Tracing Machine Unlearning to Knowledge-Tracing for Foundation Models
by: Tan, Yuwen, et al.
Published: (2025)
by: Tan, Yuwen, et al.
Published: (2025)
CAPF: Guiding Search-Agent Rollouts with Credit-Attenuated Privileged Feedback
by: Chen, Bin, et al.
Published: (2026)
by: Chen, Bin, et al.
Published: (2026)
Context-Enriched Contrastive Loss: Enhancing Presentation of Inherent Sample Connections in Contrastive Learning Framework
by: Deng, Haojin, et al.
Published: (2025)
by: Deng, Haojin, et al.
Published: (2025)
Similar Items
-
360-LLaMA-Factory: Plug & Play Sequence Parallelism for Long Post-Training
by: Zou, Haosheng, et al.
Published: (2025) -
Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
by: Wen, Liang, et al.
Published: (2025) -
Light-IF: Endowing LLMs with Generalizable Reasoning via Preview and Self-Checking for Complex Instruction Following
by: Wang, Chenyang, et al.
Published: (2025) -
The role of lipoylation in mitochondrial adaptation to methionine restriction
by: Jingyuan Xue, et al.
Published: (2024) -
Phospholipid Biosynthesis: An Unforeseen Modulator of Nuclear Metabolism
by: Hong Qiu, et al.
Published: (2025)