SAPO: Step-Aligned Policy Optimization for Reasoning-Based Generative Recommendation
Fuente:
arXiv
Saved in:
| Main Authors: | Zheng, Zaiyi, Min, Guanghui, Zhu, Yaochen, Wu, Liang, Hong, Liangjie, Chen, Chen, Li, Jundong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Knowledge Editing for Large Language Models: A Survey
by: Wang, Song, et al.
Published: (2023)
by: Wang, Song, et al.
Published: (2023)
SemCoT: Accelerating Chain-of-Thought Reasoning through Semantically-Aligned Implicit Tokens
by: He, Yinhan, et al.
Published: (2025)
by: He, Yinhan, et al.
Published: (2025)
CausalFlip: A Benchmark for LLM Causal Judgment Beyond Semantic Matching
by: Wang, Yuzhe, et al.
Published: (2026)
by: Wang, Yuzhe, et al.
Published: (2026)
Collaborative Large Language Model for Recommender Systems
by: Zhu, Yaochen, et al.
Published: (2023)
by: Zhu, Yaochen, et al.
Published: (2023)
Generative Risk Minimization for Out-of-Distribution Generalization on Graphs
by: Wang, Song, et al.
Published: (2025)
by: Wang, Song, et al.
Published: (2025)
LLM-Enhanced User-Item Interactions: Leveraging Edge Information for Optimized Recommendations
by: Wang, Xinyuan, et al.
Published: (2024)
by: Wang, Xinyuan, et al.
Published: (2024)
Knowledge Graph-Enhanced Large Language Models via Path Selection
by: Liu, Haochen, et al.
Published: (2024)
by: Liu, Haochen, et al.
Published: (2024)
KG-CF: Knowledge Graph Completion with Context Filtering under the Guidance of Large Language Models
by: Zheng, Zaiyi, et al.
Published: (2025)
by: Zheng, Zaiyi, et al.
Published: (2025)
ACPO: Adaptive Curriculum Policy Optimization for Aligning Vision-Language Models in Complex Reasoning
by: Wang, Yunhao, et al.
Published: (2025)
by: Wang, Yunhao, et al.
Published: (2025)
MA-SAPO: Multi-Agent Reasoning for Score-Aware Prompt Optimization
by: Seo, Wonduk, et al.
Published: (2025)
by: Seo, Wonduk, et al.
Published: (2025)
Segment-Aligned Policy Optimization for Multi-Modal Reasoning
by: Gao, Lei, et al.
Published: (2026)
by: Gao, Lei, et al.
Published: (2026)
Optimus-3: Dual-Router Aligned Mixture-of-Experts Agent with Dual-Granularity Reasoning-Aware Policy Optimization
by: Li, Zaijing, et al.
Published: (2025)
by: Li, Zaijing, et al.
Published: (2025)
Step Guided Reasoning: Improving Mathematical Reasoning using Guidance Generation and Step Reasoning
by: Cao, Lang, et al.
Published: (2024)
by: Cao, Lang, et al.
Published: (2024)
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization
by: Liu, Jiacai, et al.
Published: (2024)
by: Liu, Jiacai, et al.
Published: (2024)
Tool-Augmented Policy Optimization: Synergizing Reasoning and Adaptive Tool Use with Reinforcement Learning
by: Wu, Wenxun, et al.
Published: (2025)
by: Wu, Wenxun, et al.
Published: (2025)
Step-level Value Preference Optimization for Mathematical Reasoning
by: Chen, Guoxin, et al.
Published: (2024)
by: Chen, Guoxin, et al.
Published: (2024)
SE-Agent: Self-Evolution Trajectory Optimization in Multi-Step Reasoning with LLM-Based Agents
by: Lin, Jiaye, et al.
Published: (2025)
by: Lin, Jiaye, et al.
Published: (2025)
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning
by: Chen, Haoxuan, et al.
Published: (2026)
by: Chen, Haoxuan, et al.
Published: (2026)
Graph Prompting for Graph Learning Models: Recent Advances and Future Directions
by: Fu, Xingbo, et al.
Published: (2025)
by: Fu, Xingbo, et al.
Published: (2025)
One-Step Generative Policies with Q-Learning: A Reformulation of MeanFlow
by: Wang, Zeyuan, et al.
Published: (2025)
by: Wang, Zeyuan, et al.
Published: (2025)
Step-KTO: Optimizing Mathematical Reasoning through Stepwise Binary Feedback
by: Lin, Yen-Ting, et al.
Published: (2025)
by: Lin, Yen-Ting, et al.
Published: (2025)
HybridFlow: A Two-Step Generative Policy for Robotic Manipulation
by: Dong, Zhenchen, et al.
Published: (2026)
by: Dong, Zhenchen, et al.
Published: (2026)
IB-GRPO: Aligning LLM-based Learning Path Recommendation with Educational Objectives via Indicator-Based Group Relative Policy Optimization
by: Wang, Shuai, et al.
Published: (2026)
by: Wang, Shuai, et al.
Published: (2026)
Process Supervision-Guided Policy Optimization for Code Generation
by: Dai, Ning, et al.
Published: (2024)
by: Dai, Ning, et al.
Published: (2024)
Step-GRPO: Internalizing Dynamic Early Exit for Efficient Reasoning
by: Chen, Benteng, et al.
Published: (2026)
by: Chen, Benteng, et al.
Published: (2026)
From Single-Step Edit Response to Multi-Step Molecular Optimization
by: Rao, Haojie, et al.
Published: (2026)
by: Rao, Haojie, et al.
Published: (2026)
Exploring the Role of Reasoning Structures for Constructing Proofs in Multi-Step Natural Language Reasoning with Large Language Models
by: Zheng, Zi'ou, et al.
Published: (2024)
by: Zheng, Zi'ou, et al.
Published: (2024)
One-Step Flow Policy: Self-Distillation for Fast Visuomotor Policies
by: Li, Shaolong, et al.
Published: (2026)
by: Li, Shaolong, et al.
Published: (2026)
DisenReason: Behavior Disentanglement and Latent Reasoning for Shared-Account Sequential Recommendation
by: Cheng, Jiawei, et al.
Published: (2026)
by: Cheng, Jiawei, et al.
Published: (2026)
Aligning Deep Implicit Preferences by Learning to Reason Defensively
by: Li, Peiming, et al.
Published: (2025)
by: Li, Peiming, et al.
Published: (2025)
StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization
by: Wang, Ziliang, et al.
Published: (2025)
by: Wang, Ziliang, et al.
Published: (2025)
Iterative Semantic Reasoning from Individual to Group Interests for Generative Recommendation with LLMs
by: Zhu, Xiaofei, et al.
Published: (2026)
by: Zhu, Xiaofei, et al.
Published: (2026)
Enhancing Multi-Step Reasoning Abilities of Language Models through Direct Q-Function Optimization
by: Ji, Kaixuan, et al.
Published: (2024)
by: Ji, Kaixuan, et al.
Published: (2024)
Towards Resilient and Autonomous Networks: A BlueSky Vision on AI-Native 6G
by: Wu, Liang, et al.
Published: (2026)
by: Wu, Liang, et al.
Published: (2026)
Learning to Generate Formally Verifiable Step-by-Step Logic Reasoning via Structured Formal Intermediaries
by: Chen, Luoxin, et al.
Published: (2026)
by: Chen, Luoxin, et al.
Published: (2026)
Stochastic MeanFlow Policies: One-Step Generative Control with Entropic Mirror Descent
by: Wang, Zeyuan, et al.
Published: (2026)
by: Wang, Zeyuan, et al.
Published: (2026)
AlignGroup: Learning and Aligning Group Consensus with Member Preferences for Group Recommendation
by: Xu, Jinfeng, et al.
Published: (2024)
by: Xu, Jinfeng, et al.
Published: (2024)
Think before Recommendation: Autonomous Reasoning-enhanced Recommender
by: Kong, Xiaoyu, et al.
Published: (2025)
by: Kong, Xiaoyu, et al.
Published: (2025)
DRPO: Efficient Reasoning via Decoupled Reward Policy Optimization
by: Li, Gang, et al.
Published: (2025)
by: Li, Gang, et al.
Published: (2025)
Act2Goal: From World Model To General Goal-conditioned Policy
by: Zhou, Pengfei, et al.
Published: (2025)
by: Zhou, Pengfei, et al.
Published: (2025)
Similar Items
-
Knowledge Editing for Large Language Models: A Survey
by: Wang, Song, et al.
Published: (2023) -
SemCoT: Accelerating Chain-of-Thought Reasoning through Semantically-Aligned Implicit Tokens
by: He, Yinhan, et al.
Published: (2025) -
CausalFlip: A Benchmark for LLM Causal Judgment Beyond Semantic Matching
by: Wang, Yuzhe, et al.
Published: (2026) -
Collaborative Large Language Model for Recommender Systems
by: Zhu, Yaochen, et al.
Published: (2023) -
Generative Risk Minimization for Out-of-Distribution Generalization on Graphs
by: Wang, Song, et al.
Published: (2025)