LLM Agent Meets Agentic AI: Can LLM Agents Simulate Customers to Evaluate Agentic-AI-based Shopping Assistants?

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sun, Lu, Fu, Shihan, Yao, Bingsheng, Lu, Yuxuan, Li, Wenbo, Gu, Hansu, Gesi, Jiri, Huang, Jing, Luo, Chen, Wang, Dakuo
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916970742415360
author Sun, Lu
Fu, Shihan
Yao, Bingsheng
Lu, Yuxuan
Li, Wenbo
Gu, Hansu
Gesi, Jiri
Huang, Jing
Luo, Chen
Wang, Dakuo
author_facet Sun, Lu
Fu, Shihan
Yao, Bingsheng
Lu, Yuxuan
Li, Wenbo
Gu, Hansu
Gesi, Jiri
Huang, Jing
Luo, Chen
Wang, Dakuo
contents Agentic AI is emerging, capable of executing tasks through natural language, such as Copilot for coding or Amazon Rufus for shopping. Evaluating these systems is challenging, as their rapid evolution outpaces traditional human evaluation. Researchers have proposed LLM Agents to simulate participants as digital twins, but it remains unclear to what extent a digital twin can represent a specific customer in multi-turn interaction with an agentic AI system. In this paper, we recruited 40 human participants to shop with Amazon Rufus, collected their personas, interaction traces, and UX feedback, and then created digital twins to repeat the task. Pairwise comparison of human and digital-twin traces shows that while agents often explored more diverse choices, their action patterns aligned with humans and yielded similar design feedback. This study is the first to quantify how closely LLM agents can mirror human multi-turn interaction with an agentic AI system, highlighting their potential for scalable evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2509_21501
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LLM Agent Meets Agentic AI: Can LLM Agents Simulate Customers to Evaluate Agentic-AI-based Shopping Assistants?
Sun, Lu
Fu, Shihan
Yao, Bingsheng
Lu, Yuxuan
Li, Wenbo
Gu, Hansu
Gesi, Jiri
Huang, Jing
Luo, Chen
Wang, Dakuo
Human-Computer Interaction
Computation and Language
Agentic AI is emerging, capable of executing tasks through natural language, such as Copilot for coding or Amazon Rufus for shopping. Evaluating these systems is challenging, as their rapid evolution outpaces traditional human evaluation. Researchers have proposed LLM Agents to simulate participants as digital twins, but it remains unclear to what extent a digital twin can represent a specific customer in multi-turn interaction with an agentic AI system. In this paper, we recruited 40 human participants to shop with Amazon Rufus, collected their personas, interaction traces, and UX feedback, and then created digital twins to repeat the task. Pairwise comparison of human and digital-twin traces shows that while agents often explored more diverse choices, their action patterns aligned with humans and yielded similar design feedback. This study is the first to quantify how closely LLM agents can mirror human multi-turn interaction with an agentic AI system, highlighting their potential for scalable evaluation.
title LLM Agent Meets Agentic AI: Can LLM Agents Simulate Customers to Evaluate Agentic-AI-based Shopping Assistants?
topic Human-Computer Interaction
Computation and Language
url https://arxiv.org/abs/2509.21501