AgentRecBench: Benchmarking LLM Agent-based Personalized Recommender Systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shang, Yu, Liu, Peijie, Yan, Yuwei, Wu, Zijing, Sheng, Leheng, Yu, Yuanqing, Jiang, Chumeng, Zhang, An, Xu, Fengli, Wang, Yu, Zhang, Min, Li, Yong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908382602985472
author Shang, Yu
Liu, Peijie
Yan, Yuwei
Wu, Zijing
Sheng, Leheng
Yu, Yuanqing
Jiang, Chumeng
Zhang, An
Xu, Fengli
Wang, Yu
Zhang, Min
Li, Yong
author_facet Shang, Yu
Liu, Peijie
Yan, Yuwei
Wu, Zijing
Sheng, Leheng
Yu, Yuanqing
Jiang, Chumeng
Zhang, An
Xu, Fengli
Wang, Yu
Zhang, Min
Li, Yong
contents The emergence of agentic recommender systems powered by Large Language Models (LLMs) represents a paradigm shift in personalized recommendations, leveraging LLMs' advanced reasoning and role-playing capabilities to enable autonomous, adaptive decision-making. Unlike traditional recommendation approaches, agentic recommender systems can dynamically gather and interpret user-item interactions from complex environments, generating robust recommendation strategies that generalize across diverse scenarios. However, the field currently lacks standardized evaluation protocols to systematically assess these methods. To address this critical gap, we propose: (1) an interactive textual recommendation simulator incorporating rich user and item metadata and three typical evaluation scenarios (classic, evolving-interest, and cold-start recommendation tasks); (2) a unified modular framework for developing and studying agentic recommender systems; and (3) the first comprehensive benchmark comparing 10 classical and agentic recommendation methods. Our findings demonstrate the superiority of agentic systems and establish actionable design guidelines for their core components. The benchmark environment has been rigorously validated through an open challenge and remains publicly available with a continuously maintained leaderboard~\footnote[2]{https://tsinghua-fib-lab.github.io/AgentSocietyChallenge/pages/overview.html}, fostering ongoing community engagement and reproducible research. The benchmark is available at: \hyperlink{https://huggingface.co/datasets/SGJQovo/AgentRecBench}{https://huggingface.co/datasets/SGJQovo/AgentRecBench}.
format Preprint
id arxiv_https___arxiv_org_abs_2505_19623
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AgentRecBench: Benchmarking LLM Agent-based Personalized Recommender Systems
Shang, Yu
Liu, Peijie
Yan, Yuwei
Wu, Zijing
Sheng, Leheng
Yu, Yuanqing
Jiang, Chumeng
Zhang, An
Xu, Fengli
Wang, Yu
Zhang, Min
Li, Yong
Information Retrieval
Artificial Intelligence
The emergence of agentic recommender systems powered by Large Language Models (LLMs) represents a paradigm shift in personalized recommendations, leveraging LLMs' advanced reasoning and role-playing capabilities to enable autonomous, adaptive decision-making. Unlike traditional recommendation approaches, agentic recommender systems can dynamically gather and interpret user-item interactions from complex environments, generating robust recommendation strategies that generalize across diverse scenarios. However, the field currently lacks standardized evaluation protocols to systematically assess these methods. To address this critical gap, we propose: (1) an interactive textual recommendation simulator incorporating rich user and item metadata and three typical evaluation scenarios (classic, evolving-interest, and cold-start recommendation tasks); (2) a unified modular framework for developing and studying agentic recommender systems; and (3) the first comprehensive benchmark comparing 10 classical and agentic recommendation methods. Our findings demonstrate the superiority of agentic systems and establish actionable design guidelines for their core components. The benchmark environment has been rigorously validated through an open challenge and remains publicly available with a continuously maintained leaderboard~\footnote[2]{https://tsinghua-fib-lab.github.io/AgentSocietyChallenge/pages/overview.html}, fostering ongoing community engagement and reproducible research. The benchmark is available at: \hyperlink{https://huggingface.co/datasets/SGJQovo/AgentRecBench}{https://huggingface.co/datasets/SGJQovo/AgentRecBench}.
title AgentRecBench: Benchmarking LLM Agent-based Personalized Recommender Systems
topic Information Retrieval
Artificial Intelligence
url https://arxiv.org/abs/2505.19623