Clipping-Free Policy Optimization for Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Çağatan, Ömer Veysel, Akgün, Barış, Şahin, Gözde Gül, Zhao, Xuandong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Uncovering RL Integration in SSL Loss: Objective-Specific Implications for Data-Efficient RL
by: Çağatan, Ömer Veysel, et al.
Published: (2024)
by: Çağatan, Ömer Veysel, et al.
Published: (2024)
Failure Modes of Maximum Entropy RLHF
by: Çağatan, Ömer Veysel, et al.
Published: (2025)
by: Çağatan, Ömer Veysel, et al.
Published: (2025)
Higher Resolution, Better Generalization: Unlocking Visual Scaling in Deep Reinforcement Learning
by: Trumpp, Raphael, et al.
Published: (2026)
by: Trumpp, Raphael, et al.
Published: (2026)
SigCLR: Sigmoid Contrastive Learning of Visual Representations
by: Çağatan, Ömer Veysel
Published: (2024)
by: Çağatan, Ömer Veysel
Published: (2024)
UNSEE: Unsupervised Non-contrastive Sentence Embeddings
by: Çağatan, Ömer Veysel
Published: (2024)
by: Çağatan, Ömer Veysel
Published: (2024)
Linguistically-Informed Multilingual Instruction Tuning: Is There an Optimal Set of Languages to Tune?
by: Soykan, Gürkan, et al.
Published: (2024)
by: Soykan, Gürkan, et al.
Published: (2024)
Scalable Best-of-N Selection for Large Language Models via Self-Certainty
by: Kang, Zhewei, et al.
Published: (2025)
by: Kang, Zhewei, et al.
Published: (2025)
DCPO: Dynamic Clipping Policy Optimization
by: Yang, Shihui, et al.
Published: (2025)
by: Yang, Shihui, et al.
Published: (2025)
Clip-Low Increases Entropy and Clip-High Decreases Entropy in Reinforcement Learning of Large Language Models
by: Park, Jaesung R., et al.
Published: (2025)
by: Park, Jaesung R., et al.
Published: (2025)
Adversarial Robustness of Discriminative Self-Supervised Learning in Vision
by: Çağatan, Ömer Veysel, et al.
Published: (2025)
by: Çağatan, Ömer Veysel, et al.
Published: (2025)
Offline Policy Learning with Weight Clipping and Heaviside Composite Optimization
by: Liu, Jingren, et al.
Published: (2026)
by: Liu, Jingren, et al.
Published: (2026)
Inpainting-Guided Policy Optimization for Diffusion Large Language Models
by: Zhao, Siyan, et al.
Published: (2025)
by: Zhao, Siyan, et al.
Published: (2025)
Large Language Models as Generalist Policies for Network Optimization
by: Wu, Duo, et al.
Published: (2025)
by: Wu, Duo, et al.
Published: (2025)
Benchmarking Procedural Language Understanding for Low-Resource Languages: A Case Study on Turkish
by: Uzunoglu, Arda, et al.
Published: (2023)
by: Uzunoglu, Arda, et al.
Published: (2023)
To Clip or not to Clip: the Dynamics of SGD with Gradient Clipping in High-Dimensions
by: Marshall, Noah, et al.
Published: (2024)
by: Marshall, Noah, et al.
Published: (2024)
AGGC: Adaptive Group Gradient Clipping for Stabilizing Large Language Model Training
by: Li, Zhiyuan, et al.
Published: (2026)
by: Li, Zhiyuan, et al.
Published: (2026)
Diversity-Aware Policy Optimization for Large Language Model Reasoning
by: Yao, Jian, et al.
Published: (2025)
by: Yao, Jian, et al.
Published: (2025)
Graph-attention-based Casual Discovery with Trust Region-navigated Clipping Policy Optimization
by: Liu, Shixuan, et al.
Published: (2024)
by: Liu, Shixuan, et al.
Published: (2024)
Machine Bullshit: Characterizing the Emergent Disregard for Truth in Large Language Models
by: Liang, Kaiqu, et al.
Published: (2025)
by: Liang, Kaiqu, et al.
Published: (2025)
SoftAdaClip: A Smooth Clipping Strategy for Fair and Private Model Training
by: Soleymani, Dorsa, et al.
Published: (2025)
by: Soleymani, Dorsa, et al.
Published: (2025)
Group Causal Policy Optimization for Post-Training Large Language Models
by: Gu, Ziyin, et al.
Published: (2025)
by: Gu, Ziyin, et al.
Published: (2025)
CE-GPPO: Coordinating Entropy via Gradient-Preserving Clipping Policy Optimization in Reinforcement Learning
by: Su, Zhenpeng, et al.
Published: (2025)
by: Su, Zhenpeng, et al.
Published: (2025)
Robust Stochastic Optimization via Gradient Quantile Clipping
by: Merad, Ibrahim, et al.
Published: (2023)
by: Merad, Ibrahim, et al.
Published: (2023)
BAPO: Stabilizing Off-Policy Reinforcement Learning for LLMs via Balanced Policy Optimization with Adaptive Clipping
by: Xi, Zhiheng, et al.
Published: (2025)
by: Xi, Zhiheng, et al.
Published: (2025)
DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards
by: Hu, Haoyu, et al.
Published: (2026)
by: Hu, Haoyu, et al.
Published: (2026)
Klear-Reasoner: Advancing Reasoning Capability via Gradient-Preserving Clipping Policy Optimization
by: Su, Zhenpeng, et al.
Published: (2025)
by: Su, Zhenpeng, et al.
Published: (2025)
LEPO: Latent Reasoning Policy Optimization for Large Language Models
by: Zhou, Yuyan, et al.
Published: (2026)
by: Zhou, Yuyan, et al.
Published: (2026)
Adaptive Sampling and Clipping for Private Worst-Case Group Optimization
by: Cairney-Leeming, Max, et al.
Published: (2026)
by: Cairney-Leeming, Max, et al.
Published: (2026)
An Undetectable Watermark for Generative Image Models
by: Gunn, Sam, et al.
Published: (2024)
by: Gunn, Sam, et al.
Published: (2024)
BlendServe: Optimizing Offline Inference for Auto-regressive Large Models with Resource-aware Batching
by: Zhao, Yilong, et al.
Published: (2024)
by: Zhao, Yilong, et al.
Published: (2024)
DE-COP: Detecting Copyrighted Content in Language Models Training Data
by: Duarte, André V., et al.
Published: (2024)
by: Duarte, André V., et al.
Published: (2024)
Combining Large Language Models and Gradient-Free Optimization for Automatic Control Policy Synthesis
by: Bosio, Carlo, et al.
Published: (2025)
by: Bosio, Carlo, et al.
Published: (2025)
Robust Inference Methods for Latent Group Panel Models under Possible Group Non-Separation
by: Akgun, Oguzhan, et al.
Published: (2025)
by: Akgun, Oguzhan, et al.
Published: (2025)
Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs
by: Cai, Will, et al.
Published: (2025)
by: Cai, Will, et al.
Published: (2025)
Model-Agnostic Policy Explanations with Large Language Models
by: Xi-Jia, Zhang, et al.
Published: (2025)
by: Xi-Jia, Zhang, et al.
Published: (2025)
GVPO: Group Variance Policy Optimization for Large Language Model Post-Training
by: Zhang, Kaichen, et al.
Published: (2025)
by: Zhang, Kaichen, et al.
Published: (2025)
Words in Motion: Extracting Interpretable Control Vectors for Motion Transformers
by: Tas, Omer Sahin, et al.
Published: (2024)
by: Tas, Omer Sahin, et al.
Published: (2024)
CoGen: Learning from Feedback with Coupled Comprehension and Generation
by: Gul, Mustafa Omer, et al.
Published: (2024)
by: Gul, Mustafa Omer, et al.
Published: (2024)
LFPO: Likelihood-Free Policy Optimization for Masked Diffusion Models
by: Wei, Chenxing, et al.
Published: (2026)
by: Wei, Chenxing, et al.
Published: (2026)
PPO-Clip Attains Global Optimality: Towards Deeper Understandings of Clipping
by: Huang, Nai-Chieh, et al.
Published: (2023)
by: Huang, Nai-Chieh, et al.
Published: (2023)
Similar Items
-
Uncovering RL Integration in SSL Loss: Objective-Specific Implications for Data-Efficient RL
by: Çağatan, Ömer Veysel, et al.
Published: (2024) -
Failure Modes of Maximum Entropy RLHF
by: Çağatan, Ömer Veysel, et al.
Published: (2025) -
Higher Resolution, Better Generalization: Unlocking Visual Scaling in Deep Reinforcement Learning
by: Trumpp, Raphael, et al.
Published: (2026) -
SigCLR: Sigmoid Contrastive Learning of Visual Representations
by: Çağatan, Ömer Veysel
Published: (2024) -
UNSEE: Unsupervised Non-contrastive Sentence Embeddings
by: Çağatan, Ömer Veysel
Published: (2024)