Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Su, Xuerui, Xie, Shufang, Liu, Guoqing, Xia, Yingce, Luo, Renqian, Jin, Peiran, Ma, Zhiming, Wang, Yue, Wang, Zun, Liu, Yuting |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DGRO: Enhancing LLM Reasoning via Exploration-Exploitation Control and Reward Variance Management
by: Su, Xuerui, et al.
Published: (2025)
by: Su, Xuerui, et al.
Published: (2025)
HybriDNA: A Hybrid Transformer-Mamba2 Long-Range DNA Language Model
by: Ma, Mingqian, et al.
Published: (2025)
by: Ma, Mingqian, et al.
Published: (2025)
OThink-MR1: Stimulating multimodal generalized reasoning capabilities via dynamic reinforcement learning
by: Liu, Zhiyuan, et al.
Published: (2025)
by: Liu, Zhiyuan, et al.
Published: (2025)
MolChord: Structure-Sequence Alignment for Protein-Guided Drug Design
by: Zhang, Wei, et al.
Published: (2025)
by: Zhang, Wei, et al.
Published: (2025)
Reveal the Mystery of DPO: The Connection between DPO and RL Algorithms
by: Su, Xuerui, et al.
Published: (2025)
by: Su, Xuerui, et al.
Published: (2025)
Retrosynthetic Planning with Dual Value Networks
by: Liu, Guoqing, et al.
Published: (2023)
by: Liu, Guoqing, et al.
Published: (2023)
SE3Set: Harnessing equivariant hypergraph neural networks for molecular representation learning
by: Wu, Hongfei, et al.
Published: (2024)
by: Wu, Hongfei, et al.
Published: (2024)
Deciphering Scientific Reasoning Steps from Outcome Data for Molecule Optimization
by: Liu, Zequn, et al.
Published: (2026)
by: Liu, Zequn, et al.
Published: (2026)
ForeSci: Evaluating LLM Agents for Forward-Looking AI Research Judgment
by: Tian, Qiuyu, et al.
Published: (2026)
by: Tian, Qiuyu, et al.
Published: (2026)
What is the objective of reasoning with reinforcement learning?
by: Davis, Damek, et al.
Published: (2025)
by: Davis, Damek, et al.
Published: (2025)
Exploiting Pre-trained Models for Drug Target Affinity Prediction with Nearest Neighbors
by: Pei, Qizhi, et al.
Published: (2024)
by: Pei, Qizhi, et al.
Published: (2024)
CoPA: Benchmarking Personalized Question Answering with Data-Informed Cognitive Factors
by: Su, Hang, et al.
Published: (2026)
by: Su, Hang, et al.
Published: (2026)
Damage constitutive modeling of microfiber‐reinforced recycled aggregate concrete under cyclic loads
by: Changqing Wang, et al.
Published: (2025)
by: Changqing Wang, et al.
Published: (2025)
Trust Region Masking for Long-Horizon LLM Reinforcement Learning
by: Li, Yingru, et al.
Published: (2025)
by: Li, Yingru, et al.
Published: (2025)
TeamTR: Trust-Region Fine-Tuning for Multi-Agent LLM Coordination
by: Xie, Yi, et al.
Published: (2026)
by: Xie, Yi, et al.
Published: (2026)
Adaptive trajectory-constrained exploration strategy for deep reinforcement learning
by: Wang, Guojian, et al.
Published: (2023)
by: Wang, Guojian, et al.
Published: (2023)
FABind: Fast and Accurate Protein-Ligand Binding
by: Pei, Qizhi, et al.
Published: (2023)
by: Pei, Qizhi, et al.
Published: (2023)
TRE: Encouraging Exploration in the Trust Region
by: Huang, Chao, et al.
Published: (2026)
by: Huang, Chao, et al.
Published: (2026)
Research on reinforcement learning based warehouse robot navigation algorithm in complex warehouse layout
by: Li, Keqin, et al.
Published: (2024)
by: Li, Keqin, et al.
Published: (2024)
Trust-Region Adaptive Policy Optimization
by: Su, Mingyu, et al.
Published: (2025)
by: Su, Mingyu, et al.
Published: (2025)
Distributed scalable coupled policy algorithm for networked multi-agent reinforcement learning
by: Dai, Pengcheng, et al.
Published: (2025)
by: Dai, Pengcheng, et al.
Published: (2025)
Nature Language Model: Deciphering the Language of Nature for Scientific Discovery
by: Xia, Yingce, et al.
Published: (2025)
by: Xia, Yingce, et al.
Published: (2025)
Distributed primal-dual algorithm for constrained multi-agent reinforcement learning under coupled policies
by: Dai, Pengcheng, et al.
Published: (2025)
by: Dai, Pengcheng, et al.
Published: (2025)
Training-free retrieval-augmented generation with reinforced reasoning for flood damage nowcasting
by: Huang, Lipai, et al.
Published: (2026)
by: Huang, Lipai, et al.
Published: (2026)
Quantum algorithms for equational reasoning
by: Rattacaso, Davide, et al.
Published: (2025)
by: Rattacaso, Davide, et al.
Published: (2025)
Token-level Direct Preference Optimization
by: Zeng, Yongcheng, et al.
Published: (2024)
by: Zeng, Yongcheng, et al.
Published: (2024)
Asymmetric Proximal Policy Optimization: mini-critics boost LLM reasoning
by: Liu, Jiashun, et al.
Published: (2025)
by: Liu, Jiashun, et al.
Published: (2025)
Automatic damage detection and segmentation using deep learning algorithms in reinforced concrete structure inspections
by: Jiehui Wang, et al.
Published: (2024)
by: Jiehui Wang, et al.
Published: (2024)
Bellman operator convergence enhancements in reinforcement learning algorithms
by: Kadurha, David Krame, et al.
Published: (2025)
by: Kadurha, David Krame, et al.
Published: (2025)
Prediction of stock prices with automated reinforced learning algorithms
by: Said Yasin, et al.
Published: (2024)
by: Said Yasin, et al.
Published: (2024)
HRM-Agent: Training a recurrent reasoning model in dynamic environments using reinforcement learning
by: Dang, Long H, et al.
Published: (2025)
by: Dang, Long H, et al.
Published: (2025)
An energy-stable phase-field model for droplet icing simulations
by: Wang, Zhihua, et al.
Published: (2024)
by: Wang, Zhihua, et al.
Published: (2024)
Does Trust Affect Behavior in a Public Health Crisis? Testing an Extended Theory of Planned Behavior Model With Trust
by: Zhiming Liu, et al.
Published: (2025)
by: Zhiming Liu, et al.
Published: (2025)
Sorrel: A simple and flexible framework for multi-agent reinforcement learning
by: Gelpí, Rebekah A., et al.
Published: (2025)
by: Gelpí, Rebekah A., et al.
Published: (2025)
Curriculum reinforcement learning with measurable task representation learning
by: Wen, Yongyan, et al.
Published: (2026)
by: Wen, Yongyan, et al.
Published: (2026)
Reframing LLM Agent Security as an Agent-Human Interaction Problem
by: Wang, Peiran, et al.
Published: (2026)
by: Wang, Peiran, et al.
Published: (2026)
Preference Diffusion for Recommendation
by: Liu, Shuo, et al.
Published: (2024)
by: Liu, Shuo, et al.
Published: (2024)
Enhancing LLM-based Recommendation with Preference Hint Discovery from Knowledge Graph
by: Zhang, Yuting, et al.
Published: (2026)
by: Zhang, Yuting, et al.
Published: (2026)
Rescue path planning for urban flood: A deep reinforcement learning–based approach
by: Xiao‐Yan Li, et al.
Published: (2024)
by: Xiao‐Yan Li, et al.
Published: (2024)
Graphene oxide‐reinforced UHMWPE fishing monofilaments: Toward high strength and toughness
by: Wenyang Zhang, et al.
Published: (2025)
by: Wenyang Zhang, et al.
Published: (2025)
Similar Items
-
DGRO: Enhancing LLM Reasoning via Exploration-Exploitation Control and Reward Variance Management
by: Su, Xuerui, et al.
Published: (2025) -
HybriDNA: A Hybrid Transformer-Mamba2 Long-Range DNA Language Model
by: Ma, Mingqian, et al.
Published: (2025) -
OThink-MR1: Stimulating multimodal generalized reasoning capabilities via dynamic reinforcement learning
by: Liu, Zhiyuan, et al.
Published: (2025) -
MolChord: Structure-Sequence Alignment for Protein-Guided Drug Design
by: Zhang, Wei, et al.
Published: (2025) -
Reveal the Mystery of DPO: The Connection between DPO and RL Algorithms
by: Su, Xuerui, et al.
Published: (2025)