Reinforced Preference Optimization for Reasoning-Augmented Recommendations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gao, Jingtong, Song, Zeyu, Lu, Chi, Li, Xiaopeng, Xu, Derong, Wang, Maolin, Jiang, Peng, Gai, Kun, Cai, Qingpeng, Zhao, Xiangyu
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913151950258176
author Gao, Jingtong
Song, Zeyu
Lu, Chi
Li, Xiaopeng
Xu, Derong
Wang, Maolin
Jiang, Peng
Gai, Kun
Cai, Qingpeng
Zhao, Xiangyu
author_facet Gao, Jingtong
Song, Zeyu
Lu, Chi
Li, Xiaopeng
Xu, Derong
Wang, Maolin
Jiang, Peng
Gai, Kun
Cai, Qingpeng
Zhao, Xiangyu
contents Recommender systems are critical for delivering personalized content across digital platforms, and recent advances in Large Language Models (LLMs) offer new opportunities to enhance them with richer world knowledge and explicit reasoning capabilities. With the help of reasoning knowledge, recommendations can better infer users' underlying intents, adapt to evolving preferences, and leverage semantic relationships for improved accuracy and interpretability. However, existing reasoning-based recommendation methods often fail to fully align the LLM's reasoning process with recommendation-specific objectives due to structural disruption during integration and difficulties in translating free-form generation into accurate item predictions. In this paper, we introduce RPORec, a reinforced preference optimization framework that unifies an LLM backbone's reasoning ability with a dedicated recommendation head (Rechead) for precise item retrieval. RPORec comprises two stages: (1) Reasoning-Augmented Recommendation Modeling, where high-quality Chain-of-Thought (CoT) reasoning is generated and used as auxiliary knowledge to guide the Rechead in learning recommendation-specific representations; and (2) Advanced Reasoning Refinement and Alignment, in which the trained Rechead produces verifiable rewards to fine-tune the LLM backbone via reinforcement learning, enhancing reasoning quality, structural consistency, and task relevance. Extensive experiments on public benchmarks and large-scale online deployments show that RPORec consistently outperforms state-of-the-art LLM-based recommendation methods, demonstrating the effectiveness of reasoning-augmented recommendation modeling in real-world systems.
format Preprint
id arxiv_https___arxiv_org_abs_2605_21967
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Reinforced Preference Optimization for Reasoning-Augmented Recommendations
Gao, Jingtong
Song, Zeyu
Lu, Chi
Li, Xiaopeng
Xu, Derong
Wang, Maolin
Jiang, Peng
Gai, Kun
Cai, Qingpeng
Zhao, Xiangyu
Information Retrieval
Recommender systems are critical for delivering personalized content across digital platforms, and recent advances in Large Language Models (LLMs) offer new opportunities to enhance them with richer world knowledge and explicit reasoning capabilities. With the help of reasoning knowledge, recommendations can better infer users' underlying intents, adapt to evolving preferences, and leverage semantic relationships for improved accuracy and interpretability. However, existing reasoning-based recommendation methods often fail to fully align the LLM's reasoning process with recommendation-specific objectives due to structural disruption during integration and difficulties in translating free-form generation into accurate item predictions. In this paper, we introduce RPORec, a reinforced preference optimization framework that unifies an LLM backbone's reasoning ability with a dedicated recommendation head (Rechead) for precise item retrieval. RPORec comprises two stages: (1) Reasoning-Augmented Recommendation Modeling, where high-quality Chain-of-Thought (CoT) reasoning is generated and used as auxiliary knowledge to guide the Rechead in learning recommendation-specific representations; and (2) Advanced Reasoning Refinement and Alignment, in which the trained Rechead produces verifiable rewards to fine-tune the LLM backbone via reinforcement learning, enhancing reasoning quality, structural consistency, and task relevance. Extensive experiments on public benchmarks and large-scale online deployments show that RPORec consistently outperforms state-of-the-art LLM-based recommendation methods, demonstrating the effectiveness of reasoning-augmented recommendation modeling in real-world systems.
title Reinforced Preference Optimization for Reasoning-Augmented Recommendations
topic Information Retrieval
url https://arxiv.org/abs/2605.21967