Beyond Alignment: Expanding Reasoning Capacity via Manifold-Reshaping Policy Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Dayu, Yang, Jiaye, Li, Weikang, Liang, Jiahui, Li, Yang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reducing Cognitive Overhead in Tool Use via Multi-Small-Agent Reinforcement Learning
by: Wang, Dayu, et al.
Published: (2025)
by: Wang, Dayu, et al.
Published: (2025)
InjectFlow: Weak Guides Strong via Orthogonal Injection for Flow Matching
by: Wang, Dayu, et al.
Published: (2026)
by: Wang, Dayu, et al.
Published: (2026)
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment
by: Li, Jiawei, et al.
Published: (2024)
by: Li, Jiawei, et al.
Published: (2024)
Student Guides Teacher: Weak-to-Strong Inference via Spectral Orthogonal Exploration
by: Wang, Dayu, et al.
Published: (2026)
by: Wang, Dayu, et al.
Published: (2026)
DRPO: Efficient Reasoning via Decoupled Reward Policy Optimization
by: Li, Gang, et al.
Published: (2025)
by: Li, Gang, et al.
Published: (2025)
ERPO: Token-Level Entropy-Regulated Policy Optimization for Large Reasoning Models
by: Yu, Song, et al.
Published: (2026)
by: Yu, Song, et al.
Published: (2026)
Reference-guided Policy Optimization for Molecular Optimization via LLM Reasoning
by: Li, Xuan, et al.
Published: (2026)
by: Li, Xuan, et al.
Published: (2026)
Policy Guided Tree Search for Enhanced LLM Reasoning
by: Li, Yang
Published: (2025)
by: Li, Yang
Published: (2025)
Reflective Preference Optimization (RPO): Enhancing On-Policy Alignment via Hint-Guided Reflection
by: Zhao, Zihui, et al.
Published: (2025)
by: Zhao, Zihui, et al.
Published: (2025)
Theoretical Analysis of Sparse Optimization with Reparameterization, Weight Decay, and Adaptive Learning Rate
by: Xu, Huangyu, et al.
Published: (2026)
by: Xu, Huangyu, et al.
Published: (2026)
Beyond State-Wise Mirror Descent: Offline Policy Optimization with Parametric Policies
by: Li, Xiang, et al.
Published: (2026)
by: Li, Xiang, et al.
Published: (2026)
Beyond Importance Sampling: Rejection-Gated Policy Optimization
by: Sun, Ziwu, et al.
Published: (2026)
by: Sun, Ziwu, et al.
Published: (2026)
Beyond KL Divergence: Policy Optimization with Flexible Bregman Divergences for LLM Reasoning
by: Yuan, Rui, et al.
Published: (2026)
by: Yuan, Rui, et al.
Published: (2026)
Segment-Aligned Policy Optimization for Multi-Modal Reasoning
by: Gao, Lei, et al.
Published: (2026)
by: Gao, Lei, et al.
Published: (2026)
ANCORA: Learning to Question via Manifold-Anchored Self-Play for Verifiable Reasoning
by: Yang, Chengcao
Published: (2026)
by: Yang, Chengcao
Published: (2026)
Unearthing Gems from Stones: Policy Optimization with Negative Sample Augmentation for LLM Reasoning
by: Yang, Zhaohui, et al.
Published: (2025)
by: Yang, Zhaohui, et al.
Published: (2025)
Reshaping Reasoning in LLMs: A Theoretical Analysis of RL Training Dynamics through Pattern Selection
by: Chen, Xingwu, et al.
Published: (2025)
by: Chen, Xingwu, et al.
Published: (2025)
Reasoning as Gradient: Scaling MLE Agents Beyond Tree Search
by: Zhang, Yifei, et al.
Published: (2026)
by: Zhang, Yifei, et al.
Published: (2026)
Exploratory Memory-Augmented LLM Agent via Hybrid On- and Off-Policy Optimization
by: Liu, Zeyuan, et al.
Published: (2026)
by: Liu, Zeyuan, et al.
Published: (2026)
IIB-LPO: Latent Policy Optimization via Iterative Information Bottleneck
by: Deng, Huilin, et al.
Published: (2026)
by: Deng, Huilin, et al.
Published: (2026)
Reasoning on the Manifold: Bidirectional Consistency for Self-Verification in Diffusion Language Models
by: Ruan, Jiaoyang, et al.
Published: (2026)
by: Ruan, Jiaoyang, et al.
Published: (2026)
SOE: Sample-Efficient Robot Policy Self-Improvement via On-Manifold Exploration
by: Jin, Yang, et al.
Published: (2025)
by: Jin, Yang, et al.
Published: (2025)
TreeRPO: Tree Relative Policy Optimization
by: Yang, Zhicheng, et al.
Published: (2025)
by: Yang, Zhicheng, et al.
Published: (2025)
Beyond One-Preference-Fits-All Alignment: Multi-Objective Direct Preference Optimization
by: Zhou, Zhanhui, et al.
Published: (2023)
by: Zhou, Zhanhui, et al.
Published: (2023)
Beyond the Boundaries of Proximal Policy Optimization
by: Tan, Charlie B., et al.
Published: (2024)
by: Tan, Charlie B., et al.
Published: (2024)
Calibration-Aware Policy Optimization for Reasoning LLMs
by: Wang, Ziqi, et al.
Published: (2026)
by: Wang, Ziqi, et al.
Published: (2026)
Flow-based Policy With Distributional Reinforcement Learning in Trajectory Optimization
by: Hao, Ruijie, et al.
Published: (2026)
by: Hao, Ruijie, et al.
Published: (2026)
TRACE: Distilling Where It Matters via Token-Routed Self On-Policy Alignment
by: Wang, Jiaxuan, et al.
Published: (2026)
by: Wang, Jiaxuan, et al.
Published: (2026)
Reasoned Safety Alignment: Ensuring Jailbreak Defense via Answer-Then-Check
by: Cao, Chentao, et al.
Published: (2025)
by: Cao, Chentao, et al.
Published: (2025)
Prune-OPD: Efficient and Reliable On-Policy Distillation for Long-Horizon Reasoning
by: Yang, Zhicheng, et al.
Published: (2026)
by: Yang, Zhicheng, et al.
Published: (2026)
Unifying Group-Relative and Self-Distillation Policy Optimization via Sample Routing
by: Li, Gengsheng, et al.
Published: (2026)
by: Li, Gengsheng, et al.
Published: (2026)
IB-GRPO: Aligning LLM-based Learning Path Recommendation with Educational Objectives via Indicator-Based Group Relative Policy Optimization
by: Wang, Shuai, et al.
Published: (2026)
by: Wang, Shuai, et al.
Published: (2026)
Exploring the Frontiers of Softmax: Provable Optimization, Applications in Diffusion Model, and Beyond
by: Cao, Yang, et al.
Published: (2024)
by: Cao, Yang, et al.
Published: (2024)
ESPO: Early-Stopping Proximal Policy Optimization
by: Li, Zihang, et al.
Published: (2026)
by: Li, Zihang, et al.
Published: (2026)
How to Allocate, How to Learn? Dynamic Rollout Allocation and Advantage Modulation for Policy Optimization
by: Fang, Yangyi, et al.
Published: (2026)
by: Fang, Yangyi, et al.
Published: (2026)
An Optimization Algorithm for Multimodal Data Alignment
by: Zhang, Wei, et al.
Published: (2025)
by: Zhang, Wei, et al.
Published: (2025)
EP-GRPO: Entropy-Progress Aligned Group Relative Policy Optimization with Implicit Process Guidance
by: Yu, Song, et al.
Published: (2026)
by: Yu, Song, et al.
Published: (2026)
Manifold Approximation leads to Robust Kernel Alignment
by: Islam, Mohammad Tariqul, et al.
Published: (2025)
by: Islam, Mohammad Tariqul, et al.
Published: (2025)
ReGeS: Reciprocal Retrieval-Generation Synergy for Conversational Recommender Systems
by: Yang, Dayu, et al.
Published: (2025)
by: Yang, Dayu, et al.
Published: (2025)
Beyond Fixed Variables: Expanding-variate Time Series Forecasting via Flat Scheme and Spatio-temporal Focal Learning
by: Ma, Minbo, et al.
Published: (2025)
by: Ma, Minbo, et al.
Published: (2025)
Similar Items
-
Reducing Cognitive Overhead in Tool Use via Multi-Small-Agent Reinforcement Learning
by: Wang, Dayu, et al.
Published: (2025) -
InjectFlow: Weak Guides Strong via Orthogonal Injection for Flow Matching
by: Wang, Dayu, et al.
Published: (2026) -
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment
by: Li, Jiawei, et al.
Published: (2024) -
Student Guides Teacher: Weak-to-Strong Inference via Spectral Orthogonal Exploration
by: Wang, Dayu, et al.
Published: (2026) -
DRPO: Efficient Reasoning via Decoupled Reward Policy Optimization
by: Li, Gang, et al.
Published: (2025)