RecRM-Bench: Benchmarking Multidimensional Reward Modeling for Agentic Recommender Systems

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zeng, Wenwen, Zhang, Jinhui, Chen, Hao, Hu, Zhaoyu, Liang, Yongqi, Chai, Jiajun, Liu, Dengcan, Liu, Zhenfeng, Yan, Shurui, Xue, Minglong, Wang, Xiaohan, Lin, Wei, Yin, Guojun
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909035581669376
author Zeng, Wenwen
Zhang, Jinhui
Chen, Hao
Hu, Zhaoyu
Liang, Yongqi
Chai, Jiajun
Liu, Dengcan
Liu, Zhenfeng
Yan, Shurui
Xue, Minglong
Wang, Xiaohan
Lin, Wei
Yin, Guojun
author_facet Zeng, Wenwen
Zhang, Jinhui
Chen, Hao
Hu, Zhaoyu
Liang, Yongqi
Chai, Jiajun
Liu, Dengcan
Liu, Zhenfeng
Yan, Shurui
Xue, Minglong
Wang, Xiaohan
Lin, Wei
Yin, Guojun
contents The integration of Large Language Model (LLM) agents is transforming recommender systems from simple query-item matching towards deeply personalized and interactive recommendations. Reinforcement Learning (RL) provides an essential framework for the optimization of these agents in recommendation tasks. However, current methodologies remain limited by a reliance on single dimensional outcome-based rewards that focus exclusively on final user interactions, overlooking critical intermediate capabilities, such as instruction following and complex intent understanding. Despite the necessity for designing multi-dimensional reward, the field lacks a standardized benchmark to facilitate this development. To bridge this gap, we introduce RecRM-Bench, the largest and most comprehensive benchmark to date for agentic recommender systems. It comprises over 1 million structured entries across four core evaluation dimensions: instruction following, factual consistency, query-item relevance, and fine-grained user behavior prediction. By supporting comprehensive assessment from syntactic compliance to complex intent grounding and preference modeling, RecRM-Bench provides a foundational dataset for training sophisticated reward models. Furthermore, we propose a systematic framework for the construction of multi-dimensional reward models and the integration of a hybrid reward function, establishing a robust foundation for developing reliable and highly capable agentic recommender systems. The complete RecRM-Bench dataset is publicly available at https://huggingface.co/datasets/wwzeng/RecRM-Bench.
format Preprint
id arxiv_https___arxiv_org_abs_2605_11874
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle RecRM-Bench: Benchmarking Multidimensional Reward Modeling for Agentic Recommender Systems
Zeng, Wenwen
Zhang, Jinhui
Chen, Hao
Hu, Zhaoyu
Liang, Yongqi
Chai, Jiajun
Liu, Dengcan
Liu, Zhenfeng
Yan, Shurui
Xue, Minglong
Wang, Xiaohan
Lin, Wei
Yin, Guojun
Information Retrieval
The integration of Large Language Model (LLM) agents is transforming recommender systems from simple query-item matching towards deeply personalized and interactive recommendations. Reinforcement Learning (RL) provides an essential framework for the optimization of these agents in recommendation tasks. However, current methodologies remain limited by a reliance on single dimensional outcome-based rewards that focus exclusively on final user interactions, overlooking critical intermediate capabilities, such as instruction following and complex intent understanding. Despite the necessity for designing multi-dimensional reward, the field lacks a standardized benchmark to facilitate this development. To bridge this gap, we introduce RecRM-Bench, the largest and most comprehensive benchmark to date for agentic recommender systems. It comprises over 1 million structured entries across four core evaluation dimensions: instruction following, factual consistency, query-item relevance, and fine-grained user behavior prediction. By supporting comprehensive assessment from syntactic compliance to complex intent grounding and preference modeling, RecRM-Bench provides a foundational dataset for training sophisticated reward models. Furthermore, we propose a systematic framework for the construction of multi-dimensional reward models and the integration of a hybrid reward function, establishing a robust foundation for developing reliable and highly capable agentic recommender systems. The complete RecRM-Bench dataset is publicly available at https://huggingface.co/datasets/wwzeng/RecRM-Bench.
title RecRM-Bench: Benchmarking Multidimensional Reward Modeling for Agentic Recommender Systems
topic Information Retrieval
url https://arxiv.org/abs/2605.11874