Towards Sample-Efficient and Stable Reinforcement Learning for LLM-based Recommendation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Ding, Hongxun, Bao, Keqin, Zhang, Jizhi, Fang, Yi, Xu, Wenxin, Feng, Fuli, He, Xiangnan
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917238680846336
author Ding, Hongxun
Bao, Keqin
Zhang, Jizhi
Fang, Yi
Xu, Wenxin
Feng, Fuli
He, Xiangnan
author_facet Ding, Hongxun
Bao, Keqin
Zhang, Jizhi
Fang, Yi
Xu, Wenxin
Feng, Fuli
He, Xiangnan
contents While Long Chain-of-Thought (Long CoT) reasoning has shown promise in Large Language Models (LLMs), its adoption for enhancing recommendation quality is growing rapidly. In this work, we critically examine this trend and argue that Long CoT is inherently ill-suited for the sequential recommendation domain. We attribute this misalignment to two primary factors: excessive inference latency and the lack of explicit cognitive reasoning patterns in user behavioral data. Driven by these observations, we propose pivoting away from the CoT structure to directly leverage its underlying mechanism: Reinforcement Learning (RL), to explore the item space. However, applying RL directly faces significant obstacles, notably low sample efficiency-where most actions fail to provide learning signals-and training instability. To overcome these limitations, we propose RISER, a novel Reinforced Item Space Exploration framework for Recommendation. RISER is designed to transform non-learnable trajectories into effective pairwise preference data for optimization. Furthermore, it incorporates specific strategies to ensure stability, including the prevention of redundant rollouts and the constraint of token-level update magnitudes. Extensive experiments on three real-world datasets show that RISER significantly outperforms competitive baselines, establishing a robust paradigm for RL-enhanced LLM recommendation. Our code will be available at https://anonymous.4open.science/r/RISER/.
format Preprint
id arxiv_https___arxiv_org_abs_2602_00632
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Towards Sample-Efficient and Stable Reinforcement Learning for LLM-based Recommendation
Ding, Hongxun
Bao, Keqin
Zhang, Jizhi
Fang, Yi
Xu, Wenxin
Feng, Fuli
He, Xiangnan
Information Retrieval
While Long Chain-of-Thought (Long CoT) reasoning has shown promise in Large Language Models (LLMs), its adoption for enhancing recommendation quality is growing rapidly. In this work, we critically examine this trend and argue that Long CoT is inherently ill-suited for the sequential recommendation domain. We attribute this misalignment to two primary factors: excessive inference latency and the lack of explicit cognitive reasoning patterns in user behavioral data. Driven by these observations, we propose pivoting away from the CoT structure to directly leverage its underlying mechanism: Reinforcement Learning (RL), to explore the item space. However, applying RL directly faces significant obstacles, notably low sample efficiency-where most actions fail to provide learning signals-and training instability. To overcome these limitations, we propose RISER, a novel Reinforced Item Space Exploration framework for Recommendation. RISER is designed to transform non-learnable trajectories into effective pairwise preference data for optimization. Furthermore, it incorporates specific strategies to ensure stability, including the prevention of redundant rollouts and the constraint of token-level update magnitudes. Extensive experiments on three real-world datasets show that RISER significantly outperforms competitive baselines, establishing a robust paradigm for RL-enhanced LLM recommendation. Our code will be available at https://anonymous.4open.science/r/RISER/.
title Towards Sample-Efficient and Stable Reinforcement Learning for LLM-based Recommendation
topic Information Retrieval
url https://arxiv.org/abs/2602.00632