ShoppingComp: Are LLMs Really Ready for Your Shopping Cart?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tou, Huaixiao, Zeng, Ying, Li, Yuemeng, Ma, Cong, Li, Muzhi, Li, Minghao, Yuan, Weijie, Zhang, He, Jia, Kai
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917256388149248
author Tou, Huaixiao
Zeng, Ying
Li, Yuemeng
Ma, Cong
Li, Muzhi
Li, Minghao
Yuan, Weijie
Zhang, He
Jia, Kai
author_facet Tou, Huaixiao
Zeng, Ying
Li, Yuemeng
Ma, Cong
Li, Muzhi
Li, Minghao
Yuan, Weijie
Zhang, He
Jia, Kai
contents We present ShoppingComp, a challenging real-world benchmark for comprehensively evaluating LLM-powered shopping agents on three core capabilities: precise product retrieval, expert-level report generation, and safety critical decision making. Unlike prior e-commerce benchmarks, ShoppingComp introduces difficult product discovery queries with many constraints, while guaranteeing open-world products and enabling easy verification of agent outputs. The benchmark comprises 145 instances and 558 scenarios, curated by 35 experts to reflect authentic shopping needs. Results reveal stark limitations of current LLMs: even state-of-the-art models achieve low performance (e.g., 17.76\% for GPT-5.2, 15.82\% for Gemini-3-Pro).Error analysis reflects limitations in core agent competencies, including information grounding in open-world environments, reliable verification of multi-constraint requirements, consistent reasoning over noisy and conflicting evidence, and risk-aware decision making. By exposing these capability gaps, ShoppingComp characterizes the trust threshold that AI systems must cross before they can be proactively trusted for reliable real-world decision making. Our code and dataset are available at https://github.com/ByteDance-BandAI/ShoppingComp.
format Preprint
id arxiv_https___arxiv_org_abs_2511_22978
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ShoppingComp: Are LLMs Really Ready for Your Shopping Cart?
Tou, Huaixiao
Zeng, Ying
Li, Yuemeng
Ma, Cong
Li, Muzhi
Li, Minghao
Yuan, Weijie
Zhang, He
Jia, Kai
Computation and Language
We present ShoppingComp, a challenging real-world benchmark for comprehensively evaluating LLM-powered shopping agents on three core capabilities: precise product retrieval, expert-level report generation, and safety critical decision making. Unlike prior e-commerce benchmarks, ShoppingComp introduces difficult product discovery queries with many constraints, while guaranteeing open-world products and enabling easy verification of agent outputs. The benchmark comprises 145 instances and 558 scenarios, curated by 35 experts to reflect authentic shopping needs. Results reveal stark limitations of current LLMs: even state-of-the-art models achieve low performance (e.g., 17.76\% for GPT-5.2, 15.82\% for Gemini-3-Pro).Error analysis reflects limitations in core agent competencies, including information grounding in open-world environments, reliable verification of multi-constraint requirements, consistent reasoning over noisy and conflicting evidence, and risk-aware decision making. By exposing these capability gaps, ShoppingComp characterizes the trust threshold that AI systems must cross before they can be proactively trusted for reliable real-world decision making. Our code and dataset are available at https://github.com/ByteDance-BandAI/ShoppingComp.
title ShoppingComp: Are LLMs Really Ready for Your Shopping Cart?
topic Computation and Language
url https://arxiv.org/abs/2511.22978