ECom-Bench: Can LLM Agent Resolve Real-World E-commerce Customer Support Issues?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Haoxin, Peng, Xianhan, Huang, Xucheng, Huang, Yizhe, Gong, Ming, Yang, Chenghan, Liu, Yang, Jiang, Ling |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MindFlow: Revolutionizing E-commerce Customer Support with Multimodal LLM Agents
von: Gong, Ming, et al.
Veröffentlicht: (2025)
von: Gong, Ming, et al.
Veröffentlicht: (2025)
MindFlow+: A Self-Evolving Agent for E-Commerce Customer Service
von: Gong, Ming, et al.
Veröffentlicht: (2025)
von: Gong, Ming, et al.
Veröffentlicht: (2025)
PerfBench: Can Agents Resolve Real-World Performance Bugs?
von: Garg, Spandan, et al.
Veröffentlicht: (2025)
von: Garg, Spandan, et al.
Veröffentlicht: (2025)
EComStage: Stage-wise and Orientation-specific Benchmarking for Large Language Models in E-commerce
von: Zhao, Kaiyan, et al.
Veröffentlicht: (2026)
von: Zhao, Kaiyan, et al.
Veröffentlicht: (2026)
MemOrb: A Plug-and-Play Verbal-Reinforcement Memory Layer for E-Commerce Customer Service
von: Huang, Yizhe, et al.
Veröffentlicht: (2025)
von: Huang, Yizhe, et al.
Veröffentlicht: (2025)
SpatialBench: Can Agents Analyze Real-World Spatial Biology Data?
von: Workman, Kenny, et al.
Veröffentlicht: (2025)
von: Workman, Kenny, et al.
Veröffentlicht: (2025)
PTCG-Bench: Can LLM Agents Master Pokémon Trading Card Game?
von: Hua, Dongdong, et al.
Veröffentlicht: (2026)
von: Hua, Dongdong, et al.
Veröffentlicht: (2026)
AD-Bench: A Real-World, Trajectory-Aware Advertising Analytics Benchmark for LLM Agents
von: Hu, Lingxiang, et al.
Veröffentlicht: (2026)
von: Hu, Lingxiang, et al.
Veröffentlicht: (2026)
Unraveling Student Anxiety in EFL Classrooms in China: The Roles of the Classroom Environment and Control–Value Appraisals
von: Quan Qian, et al.
Veröffentlicht: (2026)
von: Quan Qian, et al.
Veröffentlicht: (2026)
EcomBench: Towards Holistic Evaluation of Foundation Agents in E-commerce
von: Min, Rui, et al.
Veröffentlicht: (2025)
von: Min, Rui, et al.
Veröffentlicht: (2025)
FinSafetyBench: Evaluating LLM Safety in Real-World Financial Scenarios
von: Hou, Yutao, et al.
Veröffentlicht: (2026)
von: Hou, Yutao, et al.
Veröffentlicht: (2026)
DeliveryBench: Can Agents Earn Profit in Real World?
von: Mao, Lingjun, et al.
Veröffentlicht: (2025)
von: Mao, Lingjun, et al.
Veröffentlicht: (2025)
SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
von: Jimenez, Carlos E., et al.
Veröffentlicht: (2023)
von: Jimenez, Carlos E., et al.
Veröffentlicht: (2023)
ConsintBench: Evaluating Language Models on Real-World Consumer Intent Understanding
von: Li, Xiaozhe, et al.
Veröffentlicht: (2025)
von: Li, Xiaozhe, et al.
Veröffentlicht: (2025)
ProAgentBench: Evaluating LLM Agents for Proactive Assistance with Real-World Data
von: Tang, Yuanbo, et al.
Veröffentlicht: (2026)
von: Tang, Yuanbo, et al.
Veröffentlicht: (2026)
Is Teacher Collaboration Enough? The Essential Role of Job Crafting in Supporting Teachers' Basic Psychological Need Satisfaction
von: Xianhan Huang, et al.
Veröffentlicht: (2026)
von: Xianhan Huang, et al.
Veröffentlicht: (2026)
Characterizing the Failure Modes of LLMs in Resolving Real-World GitHub Issues
von: Jiang, Yanjie, et al.
Veröffentlicht: (2026)
von: Jiang, Yanjie, et al.
Veröffentlicht: (2026)
Xmodel-VLM: A Simple Baseline for Multimodal Vision Language Model
von: Xu, Wanting, et al.
Veröffentlicht: (2024)
von: Xu, Wanting, et al.
Veröffentlicht: (2024)
Survey of Specialized Large Language Model
von: Yang, Chenghan, et al.
Veröffentlicht: (2025)
von: Yang, Chenghan, et al.
Veröffentlicht: (2025)
Speaker Verification in Agent-Generated Conversations
von: Yang, Yizhe, et al.
Veröffentlicht: (2024)
von: Yang, Yizhe, et al.
Veröffentlicht: (2024)
A step to compute the determinant of finite semigroups not in ECom
von: Shahzamanian, M. H.
Veröffentlicht: (2024)
von: Shahzamanian, M. H.
Veröffentlicht: (2024)
Agent^2 RL-Bench: Can LLM Agents Engineer Agentic RL Post-Training?
von: Chen, Wanyi, et al.
Veröffentlicht: (2026)
von: Chen, Wanyi, et al.
Veröffentlicht: (2026)
ISO-Bench: Can Coding Agents Optimize Real-World Inference Workloads?
von: Nangia, Ayush, et al.
Veröffentlicht: (2026)
von: Nangia, Ayush, et al.
Veröffentlicht: (2026)
FinToolBench: Evaluating LLM Agents for Real-World Financial Tool Use
von: Lu, Jiaxuan, et al.
Veröffentlicht: (2026)
von: Lu, Jiaxuan, et al.
Veröffentlicht: (2026)
PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments
von: Liu, Ruoqi, et al.
Veröffentlicht: (2026)
von: Liu, Ruoqi, et al.
Veröffentlicht: (2026)
Agent-SafetyBench: Evaluating the Safety of LLM Agents
von: Zhang, Zhexin, et al.
Veröffentlicht: (2024)
von: Zhang, Zhexin, et al.
Veröffentlicht: (2024)
Can LLM Agents Simulate Multi-Turn Human Behavior? Evidence from Real Online Customer Behavior Data
von: Lu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Lu, Yuxuan, et al.
Veröffentlicht: (2025)
SessionIntentBench: A Multi-task Inter-session Intention-shift Modeling Benchmark for E-commerce Customer Behavior Understanding
von: Yang, Yuqi, et al.
Veröffentlicht: (2025)
von: Yang, Yuqi, et al.
Veröffentlicht: (2025)
End-Cloud Collaboration Framework for Advanced AI Customer Service in E-commerce
von: Teng, Liangyu, et al.
Veröffentlicht: (2024)
von: Teng, Liangyu, et al.
Veröffentlicht: (2024)
Xmodel-LM Technical Report
von: Wang, Yichuan, et al.
Veröffentlicht: (2024)
von: Wang, Yichuan, et al.
Veröffentlicht: (2024)
StockBench: Can LLM Agents Trade Stocks Profitably In Real-world Markets?
von: Chen, Yanxu, et al.
Veröffentlicht: (2025)
von: Chen, Yanxu, et al.
Veröffentlicht: (2025)
LLM-Based World Models Can Make Decisions Solely, But Rigorous Evaluations are Needed
von: Yang, Chang, et al.
Veröffentlicht: (2024)
von: Yang, Chang, et al.
Veröffentlicht: (2024)
You Are What You Bought: Generating Customer Personas for E-commerce Applications
von: Shi, Yimin, et al.
Veröffentlicht: (2025)
von: Shi, Yimin, et al.
Veröffentlicht: (2025)
HWE-Bench: Benchmarking LLM Agents on Real-World Hardware Bug Repair Tasks
von: Cui, Fan, et al.
Veröffentlicht: (2026)
von: Cui, Fan, et al.
Veröffentlicht: (2026)
DataGovBench: Benchmarking LLM Agents for Real-World Data Governance Workflows
von: Liu, Zhou, et al.
Veröffentlicht: (2025)
von: Liu, Zhou, et al.
Veröffentlicht: (2025)
BuildBench: Benchmarking LLM Agents on Compiling Real-World Open-Source Software
von: Zhang, Zehua, et al.
Veröffentlicht: (2025)
von: Zhang, Zehua, et al.
Veröffentlicht: (2025)
LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks
von: Long, Xiang, et al.
Veröffentlicht: (2026)
von: Long, Xiang, et al.
Veröffentlicht: (2026)
REI-Bench: Can Embodied Agents Understand Vague Human Instructions in Task Planning?
von: Jiang, Chenxi, et al.
Veröffentlicht: (2025)
von: Jiang, Chenxi, et al.
Veröffentlicht: (2025)
LLM Agent Meets Agentic AI: Can LLM Agents Simulate Customers to Evaluate Agentic-AI-based Shopping Assistants?
von: Sun, Lu, et al.
Veröffentlicht: (2025)
von: Sun, Lu, et al.
Veröffentlicht: (2025)
OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation
von: Hu, Xiaomeng, et al.
Veröffentlicht: (2026)
von: Hu, Xiaomeng, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
MindFlow: Revolutionizing E-commerce Customer Support with Multimodal LLM Agents
von: Gong, Ming, et al.
Veröffentlicht: (2025) -
MindFlow+: A Self-Evolving Agent for E-Commerce Customer Service
von: Gong, Ming, et al.
Veröffentlicht: (2025) -
PerfBench: Can Agents Resolve Real-World Performance Bugs?
von: Garg, Spandan, et al.
Veröffentlicht: (2025) -
EComStage: Stage-wise and Orientation-specific Benchmarking for Large Language Models in E-commerce
von: Zhao, Kaiyan, et al.
Veröffentlicht: (2026) -
MemOrb: A Plug-and-Play Verbal-Reinforcement Memory Layer for E-Commerce Customer Service
von: Huang, Yizhe, et al.
Veröffentlicht: (2025)