BandPO: Bridging Trust Regions and Ratio Clipping via Probability-Aware Bounds for LLM Reinforcement Learning
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Yuan, Wang, Bo, Gao, Yufei, Yao, Yuqian, Wang, Xinyuan, Yin, Zhangyue, Qiu, Xipeng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
It's Not You, It's Clipping: A Soft Trust-Region via Probability Smoothing for LLM RL
di: Dwyer, Madeleine, et al.
Pubblicazione: (2025)
di: Dwyer, Madeleine, et al.
Pubblicazione: (2025)
LLatrieval: LLM-Verified Retrieval for Verifiable Generation
di: Li, Xiaonan, et al.
Pubblicazione: (2023)
di: Li, Xiaonan, et al.
Pubblicazione: (2023)
RLoop: An Self-Improving Framework for Reinforcement Learning with Iterative Policy Initialization
di: Zhiyuan, Zeng, et al.
Pubblicazione: (2025)
di: Zhiyuan, Zeng, et al.
Pubblicazione: (2025)
R3-RAG: Learning Step-by-Step Reasoning and Retrieval for LLMs via Reinforcement Learning
di: Li, Yuan, et al.
Pubblicazione: (2025)
di: Li, Yuan, et al.
Pubblicazione: (2025)
Scaling of Search and Learning: A Roadmap to Reproduce o1 from Reinforcement Learning Perspective
di: Zeng, Zhiyuan, et al.
Pubblicazione: (2024)
di: Zeng, Zhiyuan, et al.
Pubblicazione: (2024)
Revisiting the Test-Time Scaling of o1-like Models: Do they Truly Possess Test-Time Scaling Capabilities?
di: Zeng, Zhiyuan, et al.
Pubblicazione: (2025)
di: Zeng, Zhiyuan, et al.
Pubblicazione: (2025)
Error Classification of Large Language Models on Math Word Problems: A Dynamically Adaptive Framework
di: Sun, Yuhong, et al.
Pubblicazione: (2025)
di: Sun, Yuhong, et al.
Pubblicazione: (2025)
ARISE: An Adaptive Resolution-Aware Metric for Test-Time Scaling Evaluation in Large Reasoning Models
di: Yin, Zhangyue, et al.
Pubblicazione: (2025)
di: Yin, Zhangyue, et al.
Pubblicazione: (2025)
Online Distributed Optimization with Clipped Stochastic Gradients: High Probability Bound of Regrets
di: Yang, Yuchen, et al.
Pubblicazione: (2024)
di: Yang, Yuchen, et al.
Pubblicazione: (2024)
Dynamic and Generalizable Process Reward Modeling
di: Yin, Zhangyue, et al.
Pubblicazione: (2025)
di: Yin, Zhangyue, et al.
Pubblicazione: (2025)
Benchmarking Hallucination in Large Language Models based on Unanswerable Math Word Problem
di: Sun, Yuhong, et al.
Pubblicazione: (2024)
di: Sun, Yuhong, et al.
Pubblicazione: (2024)
Corex: Pushing the Boundaries of Complex Reasoning through Multi-Model Collaboration
di: Sun, Qiushi, et al.
Pubblicazione: (2023)
di: Sun, Qiushi, et al.
Pubblicazione: (2023)
TRAM: Bridging Trust Regions and Sharpness Aware Minimization
di: Sherborne, Tom, et al.
Pubblicazione: (2023)
di: Sherborne, Tom, et al.
Pubblicazione: (2023)
Adaptive-Boundary-Clipping GRPO: Ensuring Bounded Ratios for Stable and Generalizable Training
di: Liu, Chi, et al.
Pubblicazione: (2026)
di: Liu, Chi, et al.
Pubblicazione: (2026)
Trust Region Masking for Long-Horizon LLM Reinforcement Learning
di: Li, Yingru, et al.
Pubblicazione: (2025)
di: Li, Yingru, et al.
Pubblicazione: (2025)
REARANK: Reasoning Re-ranking Agent via Reinforcement Learning
di: Zhang, Le, et al.
Pubblicazione: (2025)
di: Zhang, Le, et al.
Pubblicazione: (2025)
FamilyTool: A Multi-hop Personalized Tool Use Benchmark
di: Wang, Yuxin, et al.
Pubblicazione: (2025)
di: Wang, Yuxin, et al.
Pubblicazione: (2025)
VehicleWorld: A Highly Integrated Multi-Device Environment for Intelligent Vehicle Interaction
di: Yang, Jie, et al.
Pubblicazione: (2025)
di: Yang, Jie, et al.
Pubblicazione: (2025)
Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning
di: Su, Zhenpeng, et al.
Pubblicazione: (2025)
di: Su, Zhenpeng, et al.
Pubblicazione: (2025)
Differentially Private Clipped-SGD: High-Probability Convergence with Arbitrary Clipping Level
di: Khah, Saleh Vatan, et al.
Pubblicazione: (2025)
di: Khah, Saleh Vatan, et al.
Pubblicazione: (2025)
Bridging Internal Probability and Self-Consistency for Effective and Efficient LLM Reasoning
di: Zhou, Zhi, et al.
Pubblicazione: (2025)
di: Zhou, Zhi, et al.
Pubblicazione: (2025)
Rethinking Ratio-Based Trust Regions for Policy Optimization in Multi-Agent Reinforcement Learning
di: Wijesundara, Chulabhaya, et al.
Pubblicazione: (2026)
di: Wijesundara, Chulabhaya, et al.
Pubblicazione: (2026)
High Probability Complexity Bounds of Trust-Region Stochastic Sequential Quadratic Programming with Heavy-Tailed Noise
di: Fang, Yuchen, et al.
Pubblicazione: (2025)
di: Fang, Yuchen, et al.
Pubblicazione: (2025)
Bounded Ratio Reinforcement Learning
di: Ao, Yunke, et al.
Pubblicazione: (2026)
di: Ao, Yunke, et al.
Pubblicazione: (2026)
Graph-attention-based Casual Discovery with Trust Region-navigated Clipping Policy Optimization
di: Liu, Shixuan, et al.
Pubblicazione: (2024)
di: Liu, Shixuan, et al.
Pubblicazione: (2024)
A Theoretical Study on Bridging Internal Probability and Self-Consistency for LLM Reasoning
di: Zhou, Zhi, et al.
Pubblicazione: (2025)
di: Zhou, Zhi, et al.
Pubblicazione: (2025)
TeamTR: Trust-Region Fine-Tuning for Multi-Agent LLM Coordination
di: Xie, Yi, et al.
Pubblicazione: (2026)
di: Xie, Yi, et al.
Pubblicazione: (2026)
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
di: Zhang, Shiduo, et al.
Pubblicazione: (2024)
di: Zhang, Shiduo, et al.
Pubblicazione: (2024)
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping
di: Liu, Yang, et al.
Pubblicazione: (2026)
di: Liu, Yang, et al.
Pubblicazione: (2026)
High Probability Convergence of Distributed Clipped Stochastic Gradient Descent with Heavy-tailed Noise
di: Yang, Yuchen, et al.
Pubblicazione: (2025)
di: Yang, Yuchen, et al.
Pubblicazione: (2025)
Clipped SGD Algorithms for Performative Prediction: Tight Bounds for Clipping Bias and Remedies
di: Li, Qiang, et al.
Pubblicazione: (2024)
di: Li, Qiang, et al.
Pubblicazione: (2024)
GeoClip: Geometry-Aware Clipping for Differentially Private SGD
di: Gilani, Atefeh, et al.
Pubblicazione: (2025)
di: Gilani, Atefeh, et al.
Pubblicazione: (2025)
ConfClip: Confidence-Weighted and Clipped Reward for Reinforcement Learning in LLMs
di: Zhang, Bonan, et al.
Pubblicazione: (2025)
di: Zhang, Bonan, et al.
Pubblicazione: (2025)
Trust Region On-Policy Distillation
di: Xing, Xingrun, et al.
Pubblicazione: (2026)
di: Xing, Xingrun, et al.
Pubblicazione: (2026)
LLM can Achieve Self-Regulation via Hyperparameter Aware Generation
di: Wang, Siyin, et al.
Pubblicazione: (2024)
di: Wang, Siyin, et al.
Pubblicazione: (2024)
Turn Waste into Worth: Rectifying Top-$k$ Router of MoE
di: Zeng, Zhiyuan, et al.
Pubblicazione: (2024)
di: Zeng, Zhiyuan, et al.
Pubblicazione: (2024)
Are LLMs Rational Investors? A Study on Detecting and Reducing the Financial Bias in LLMs
di: Zhou, Yuhang, et al.
Pubblicazione: (2024)
di: Zhou, Yuhang, et al.
Pubblicazione: (2024)
Can AI Assistants Know What They Don't Know?
di: Cheng, Qinyuan, et al.
Pubblicazione: (2024)
di: Cheng, Qinyuan, et al.
Pubblicazione: (2024)
Unified Active Retrieval for Retrieval Augmented Generation
di: Cheng, Qinyuan, et al.
Pubblicazione: (2024)
di: Cheng, Qinyuan, et al.
Pubblicazione: (2024)
SetPO: Set-Level Policy Optimization for Diversity-Preserving LLM Reasoning
di: Li, Chenyi, et al.
Pubblicazione: (2026)
di: Li, Chenyi, et al.
Pubblicazione: (2026)
Documenti analoghi
-
It's Not You, It's Clipping: A Soft Trust-Region via Probability Smoothing for LLM RL
di: Dwyer, Madeleine, et al.
Pubblicazione: (2025) -
LLatrieval: LLM-Verified Retrieval for Verifiable Generation
di: Li, Xiaonan, et al.
Pubblicazione: (2023) -
RLoop: An Self-Improving Framework for Reinforcement Learning with Iterative Policy Initialization
di: Zhiyuan, Zeng, et al.
Pubblicazione: (2025) -
R3-RAG: Learning Step-by-Step Reasoning and Retrieval for LLMs via Reinforcement Learning
di: Li, Yuan, et al.
Pubblicazione: (2025) -
Scaling of Search and Learning: A Roadmap to Reproduce o1 from Reinforcement Learning Perspective
di: Zeng, Zhiyuan, et al.
Pubblicazione: (2024)