ShopSimulator: Evaluating and Exploring RL-Driven LLM Agent for Shopping Assistants

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Pei, Wu, Yanan, Song, Xiaoshuai, Wang, Weixun, Chen, Gengru, Li, Zhongwen, Yan, Kezhong, Deng, Ken, Liu, Qi, Zhao, Shuaibing, Xiong, Shaopan, Liu, Xuepeng, Chen, Xuefeng, Deng, Wanxi, Su, Wenbo, Zheng, Bo
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914280512684032
author Wang, Pei
Wu, Yanan
Song, Xiaoshuai
Wang, Weixun
Chen, Gengru
Li, Zhongwen
Yan, Kezhong
Deng, Ken
Liu, Qi
Zhao, Shuaibing
Xiong, Shaopan
Liu, Xuepeng
Chen, Xuefeng
Deng, Wanxi
Su, Wenbo
Zheng, Bo
author_facet Wang, Pei
Wu, Yanan
Song, Xiaoshuai
Wang, Weixun
Chen, Gengru
Li, Zhongwen
Yan, Kezhong
Deng, Ken
Liu, Qi
Zhao, Shuaibing
Xiong, Shaopan
Liu, Xuepeng
Chen, Xuefeng
Deng, Wanxi
Su, Wenbo
Zheng, Bo
contents Large language model (LLM)-based agents are increasingly deployed in e-commerce shopping. To perform thorough, user-tailored product searches, agents should interpret personal preferences, engage in multi-turn dialogues, and ultimately retrieve and discriminate among highly similar products. However, existing research has yet to provide a unified simulation environment that consistently captures all of these aspects, and always focuses solely on evaluation benchmarks without training support. In this paper, we introduce ShopSimulator, a large-scale and challenging Chinese shopping environment. Leveraging ShopSimulator, we evaluate LLMs across diverse scenarios, finding that even the best-performing models achieve less than 40% full-success rate. Error analysis reveals that agents struggle with deep search and product selection in long trajectories, fail to balance the use of personalization cues, and to effectively engage with users. Further training exploration provides practical guidance for overcoming these weaknesses, with the combination of supervised fine-tuning (SFT) and reinforcement learning (RL) yielding significant performance improvements. Code and data will be released at https://github.com/ShopAgent-Team/ShopSimulator.
format Preprint
id arxiv_https___arxiv_org_abs_2601_18225
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ShopSimulator: Evaluating and Exploring RL-Driven LLM Agent for Shopping Assistants
Wang, Pei
Wu, Yanan
Song, Xiaoshuai
Wang, Weixun
Chen, Gengru
Li, Zhongwen
Yan, Kezhong
Deng, Ken
Liu, Qi
Zhao, Shuaibing
Xiong, Shaopan
Liu, Xuepeng
Chen, Xuefeng
Deng, Wanxi
Su, Wenbo
Zheng, Bo
Artificial Intelligence
Large language model (LLM)-based agents are increasingly deployed in e-commerce shopping. To perform thorough, user-tailored product searches, agents should interpret personal preferences, engage in multi-turn dialogues, and ultimately retrieve and discriminate among highly similar products. However, existing research has yet to provide a unified simulation environment that consistently captures all of these aspects, and always focuses solely on evaluation benchmarks without training support. In this paper, we introduce ShopSimulator, a large-scale and challenging Chinese shopping environment. Leveraging ShopSimulator, we evaluate LLMs across diverse scenarios, finding that even the best-performing models achieve less than 40% full-success rate. Error analysis reveals that agents struggle with deep search and product selection in long trajectories, fail to balance the use of personalization cues, and to effectively engage with users. Further training exploration provides practical guidance for overcoming these weaknesses, with the combination of supervised fine-tuning (SFT) and reinforcement learning (RL) yielding significant performance improvements. Code and data will be released at https://github.com/ShopAgent-Team/ShopSimulator.
title ShopSimulator: Evaluating and Exploring RL-Driven LLM Agent for Shopping Assistants
topic Artificial Intelligence
url https://arxiv.org/abs/2601.18225