Think before Recommendation: Autonomous Reasoning-enhanced Recommender

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kong, Xiaoyu, Jiang, Junguang, Liu, Bin, Xu, Ziru, Zhu, Han, Xu, Jian, Zheng, Bo, Wu, Jiancan, Wang, Xiang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917045132591104
author Kong, Xiaoyu
Jiang, Junguang
Liu, Bin
Xu, Ziru
Zhu, Han
Xu, Jian
Zheng, Bo
Wu, Jiancan
Wang, Xiang
author_facet Kong, Xiaoyu
Jiang, Junguang
Liu, Bin
Xu, Ziru
Zhu, Han
Xu, Jian
Zheng, Bo
Wu, Jiancan
Wang, Xiang
contents The core task of recommender systems is to learn user preferences from historical user-item interactions. With the rapid development of large language models (LLMs), recent research has explored leveraging the reasoning capabilities of LLMs to enhance rating prediction tasks. However, existing distillation-based methods suffer from limitations such as the teacher model's insufficient recommendation capability, costly and static supervision, and superficial transfer of reasoning ability. To address these issues, this paper proposes RecZero, a reinforcement learning (RL)-based recommendation paradigm that abandons the traditional multi-model and multi-stage distillation approach. Instead, RecZero trains a single LLM through pure RL to autonomously develop reasoning capabilities for rating prediction. RecZero consists of two key components: (1) "Think-before-Recommendation" prompt construction, which employs a structured reasoning template to guide the model in step-wise analysis of user interests, item features, and user-item compatibility; and (2) rule-based reward modeling, which adopts group relative policy optimization (GRPO) to compute rewards for reasoning trajectories and optimize the LLM. Additionally, the paper explores a hybrid paradigm, RecOne, which combines supervised fine-tuning with RL, initializing the model with cold-start reasoning samples and further optimizing it with RL. Experimental results demonstrate that RecZero and RecOne significantly outperform existing baseline methods on multiple benchmark datasets, validating the superiority of the RL paradigm in achieving autonomous reasoning-enhanced recommender systems.
format Preprint
id arxiv_https___arxiv_org_abs_2510_23077
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Think before Recommendation: Autonomous Reasoning-enhanced Recommender
Kong, Xiaoyu
Jiang, Junguang
Liu, Bin
Xu, Ziru
Zhu, Han
Xu, Jian
Zheng, Bo
Wu, Jiancan
Wang, Xiang
Information Retrieval
Artificial Intelligence
The core task of recommender systems is to learn user preferences from historical user-item interactions. With the rapid development of large language models (LLMs), recent research has explored leveraging the reasoning capabilities of LLMs to enhance rating prediction tasks. However, existing distillation-based methods suffer from limitations such as the teacher model's insufficient recommendation capability, costly and static supervision, and superficial transfer of reasoning ability. To address these issues, this paper proposes RecZero, a reinforcement learning (RL)-based recommendation paradigm that abandons the traditional multi-model and multi-stage distillation approach. Instead, RecZero trains a single LLM through pure RL to autonomously develop reasoning capabilities for rating prediction. RecZero consists of two key components: (1) "Think-before-Recommendation" prompt construction, which employs a structured reasoning template to guide the model in step-wise analysis of user interests, item features, and user-item compatibility; and (2) rule-based reward modeling, which adopts group relative policy optimization (GRPO) to compute rewards for reasoning trajectories and optimize the LLM. Additionally, the paper explores a hybrid paradigm, RecOne, which combines supervised fine-tuning with RL, initializing the model with cold-start reasoning samples and further optimizing it with RL. Experimental results demonstrate that RecZero and RecOne significantly outperform existing baseline methods on multiple benchmark datasets, validating the superiority of the RL paradigm in achieving autonomous reasoning-enhanced recommender systems.
title Think before Recommendation: Autonomous Reasoning-enhanced Recommender
topic Information Retrieval
Artificial Intelligence
url https://arxiv.org/abs/2510.23077