Segment-Aligned Policy Optimization for Multi-Modal Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gao, Lei, Li, Zhuoming, Jia, Mengxi, Yuan, Jiakang, Sun, Hongbo, Sun, Hao, Li, Xuelong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Controllable Exploration in Hybrid-Policy RLVR for Multi-Modal Reasoning
von: Huang, Zhuoxu, et al.
Veröffentlicht: (2026)
von: Huang, Zhuoxu, et al.
Veröffentlicht: (2026)
Preferred-Action-Optimized Diffusion Policies for Offline Reinforcement Learning
von: Zhang, Tianle, et al.
Veröffentlicht: (2024)
von: Zhang, Tianle, et al.
Veröffentlicht: (2024)
StitchFusion: Weaving Any Visual Modalities to Enhance Multimodal Semantic Segmentation
von: Li, Bingyu, et al.
Veröffentlicht: (2024)
von: Li, Bingyu, et al.
Veröffentlicht: (2024)
Beyond Importance Sampling: Rejection-Gated Policy Optimization
von: Sun, Ziwu, et al.
Veröffentlicht: (2026)
von: Sun, Ziwu, et al.
Veröffentlicht: (2026)
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment
von: Li, Jiawei, et al.
Veröffentlicht: (2024)
von: Li, Jiawei, et al.
Veröffentlicht: (2024)
IB-GRPO: Aligning LLM-based Learning Path Recommendation with Educational Objectives via Indicator-Based Group Relative Policy Optimization
von: Wang, Shuai, et al.
Veröffentlicht: (2026)
von: Wang, Shuai, et al.
Veröffentlicht: (2026)
Multi-Modal Manipulation via Multi-Modal Policy Consensus
von: Chen, Haonan, et al.
Veröffentlicht: (2025)
von: Chen, Haonan, et al.
Veröffentlicht: (2025)
EP-GRPO: Entropy-Progress Aligned Group Relative Policy Optimization with Implicit Process Guidance
von: Yu, Song, et al.
Veröffentlicht: (2026)
von: Yu, Song, et al.
Veröffentlicht: (2026)
Infinite Video Understanding
von: Zhang, Dell, et al.
Veröffentlicht: (2025)
von: Zhang, Dell, et al.
Veröffentlicht: (2025)
Metis: Learning to Jailbreak LLMs via Self-Evolving Metacognitive Policy Optimization
von: Zhou, Huilin, et al.
Veröffentlicht: (2026)
von: Zhou, Huilin, et al.
Veröffentlicht: (2026)
Reason in Chains, Learn in Trees: Self-Rectification and Grafting for Multi-turn Agent Policy Optimization
von: Li, Yu, et al.
Veröffentlicht: (2026)
von: Li, Yu, et al.
Veröffentlicht: (2026)
Reference-guided Policy Optimization for Molecular Optimization via LLM Reasoning
von: Li, Xuan, et al.
Veröffentlicht: (2026)
von: Li, Xuan, et al.
Veröffentlicht: (2026)
Towards Robust Multi-Modal Reasoning via Model Selection
von: Liu, Xiangyan, et al.
Veröffentlicht: (2023)
von: Liu, Xiangyan, et al.
Veröffentlicht: (2023)
MMCORE: MultiModal COnnection with Representation Aligned Latent Embeddings
von: Li, Zijie, et al.
Veröffentlicht: (2026)
von: Li, Zijie, et al.
Veröffentlicht: (2026)
PVPO: Pre-Estimated Value-Based Policy Optimization for Agentic Reasoning
von: Feng, Wenfeng, et al.
Veröffentlicht: (2025)
von: Feng, Wenfeng, et al.
Veröffentlicht: (2025)
Absolute Policy Optimization
von: Zhao, Weiye, et al.
Veröffentlicht: (2023)
von: Zhao, Weiye, et al.
Veröffentlicht: (2023)
ERPO: Token-Level Entropy-Regulated Policy Optimization for Large Reasoning Models
von: Yu, Song, et al.
Veröffentlicht: (2026)
von: Yu, Song, et al.
Veröffentlicht: (2026)
Optimizing Anytime Reasoning via Budget Relative Policy Optimization
von: Qi, Penghui, et al.
Veröffentlicht: (2025)
von: Qi, Penghui, et al.
Veröffentlicht: (2025)
DRPO: Efficient Reasoning via Decoupled Reward Policy Optimization
von: Li, Gang, et al.
Veröffentlicht: (2025)
von: Li, Gang, et al.
Veröffentlicht: (2025)
Beyond Alignment: Expanding Reasoning Capacity via Manifold-Reshaping Policy Optimization
von: Wang, Dayu, et al.
Veröffentlicht: (2026)
von: Wang, Dayu, et al.
Veröffentlicht: (2026)
Beyond High-Entropy Exploration: Correctness-Aware Low-Entropy Segment-Based Advantage Shaping for Reasoning LLMs
von: Chen, Xinzhu, et al.
Veröffentlicht: (2025)
von: Chen, Xinzhu, et al.
Veröffentlicht: (2025)
Towards Learnable Anchor for Deep Multi-View Clustering
von: Wang, Bocheng, et al.
Veröffentlicht: (2025)
von: Wang, Bocheng, et al.
Veröffentlicht: (2025)
Meta-Aligner: Bidirectional Preference-Policy Optimization for Multi-Objective LLMs Alignment
von: Xu, Wenzhe, et al.
Veröffentlicht: (2026)
von: Xu, Wenzhe, et al.
Veröffentlicht: (2026)
LEPO: Latent Reasoning Policy Optimization for Large Language Models
von: Zhou, Yuyan, et al.
Veröffentlicht: (2026)
von: Zhou, Yuyan, et al.
Veröffentlicht: (2026)
SIPO: Stabilized and Improved Preference Optimization for Aligning Diffusion Models
von: Yang, Xiaomeng, et al.
Veröffentlicht: (2025)
von: Yang, Xiaomeng, et al.
Veröffentlicht: (2025)
MMICT: Boosting Multi-Modal Fine-Tuning with In-Context Examples
von: Chen, Tao, et al.
Veröffentlicht: (2023)
von: Chen, Tao, et al.
Veröffentlicht: (2023)
Dynamic Adaptive Parsing of Temporal and Cross-Variable Patterns for Network State Classification
von: Gao, Yuan, et al.
Veröffentlicht: (2025)
von: Gao, Yuan, et al.
Veröffentlicht: (2025)
Decision Flow Policy Optimization
von: Hu, Jifeng, et al.
Veröffentlicht: (2025)
von: Hu, Jifeng, et al.
Veröffentlicht: (2025)
Efficiently Aligning Draft Models via Parameter- and Data-Efficient Adaptation
von: Lin, Luxi, et al.
Veröffentlicht: (2026)
von: Lin, Luxi, et al.
Veröffentlicht: (2026)
Beyond KL Divergence: Policy Optimization with Flexible Bregman Divergences for LLM Reasoning
von: Yuan, Rui, et al.
Veröffentlicht: (2026)
von: Yuan, Rui, et al.
Veröffentlicht: (2026)
To Theoretically Understand Transformer-Based In-Context Learning for Optimizing CSMA
von: Hao, Shugang, et al.
Veröffentlicht: (2025)
von: Hao, Shugang, et al.
Veröffentlicht: (2025)
MAMMAL -- Molecular Aligned Multi-Modal Architecture and Language
von: Shoshan, Yoel, et al.
Veröffentlicht: (2024)
von: Shoshan, Yoel, et al.
Veröffentlicht: (2024)
GRPO-VPS: Enhancing Group Relative Policy Optimization with Verifiable Process Supervision for Effective Reasoning
von: Wang, Jingyi, et al.
Veröffentlicht: (2026)
von: Wang, Jingyi, et al.
Veröffentlicht: (2026)
Physics in Next-token Prediction
von: An, Hongjun, et al.
Veröffentlicht: (2024)
von: An, Hongjun, et al.
Veröffentlicht: (2024)
Regressing the Relative Future: Efficient Policy Optimization for Multi-turn RLHF
von: Gao, Zhaolin, et al.
Veröffentlicht: (2024)
von: Gao, Zhaolin, et al.
Veröffentlicht: (2024)
AutoMLGen: Navigating Fine-Grained Optimization for Coding Agents
von: Du, Shangheng, et al.
Veröffentlicht: (2025)
von: Du, Shangheng, et al.
Veröffentlicht: (2025)
Pessimistic Value Iteration for Multi-Task Data Sharing in Offline Reinforcement Learning
von: Bai, Chenjia, et al.
Veröffentlicht: (2024)
von: Bai, Chenjia, et al.
Veröffentlicht: (2024)
Calibration-Aware Policy Optimization for Reasoning LLMs
von: Wang, Ziqi, et al.
Veröffentlicht: (2026)
von: Wang, Ziqi, et al.
Veröffentlicht: (2026)
A Variance-Reduced Cubic-Regularized Newton for Policy Optimization
von: Sun, Cheng, et al.
Veröffentlicht: (2025)
von: Sun, Cheng, et al.
Veröffentlicht: (2025)
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning
von: Li, Mengqi, et al.
Veröffentlicht: (2025)
von: Li, Mengqi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Controllable Exploration in Hybrid-Policy RLVR for Multi-Modal Reasoning
von: Huang, Zhuoxu, et al.
Veröffentlicht: (2026) -
Preferred-Action-Optimized Diffusion Policies for Offline Reinforcement Learning
von: Zhang, Tianle, et al.
Veröffentlicht: (2024) -
StitchFusion: Weaving Any Visual Modalities to Enhance Multimodal Semantic Segmentation
von: Li, Bingyu, et al.
Veröffentlicht: (2024) -
Beyond Importance Sampling: Rejection-Gated Policy Optimization
von: Sun, Ziwu, et al.
Veröffentlicht: (2026) -
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment
von: Li, Jiawei, et al.
Veröffentlicht: (2024)