Benchmarking and Learning Real-World Customer Service Dialogue
Fuente:
arXiv
Saved in:
| Main Authors: | Gao, Tianhong, Shen, Jundong, Wang, Jiapeng, Shi, Bei, Ju, Ying, Yao, Junfeng, Yu, Huiyu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reinforcing Real-world Service Agents: Balancing Utility and Cost in Task-oriented Dialogue
by: Gao, Ning, et al.
Published: (2026)
by: Gao, Ning, et al.
Published: (2026)
ChatPattern: Layout Pattern Customization via Natural Language
by: Wang, Zixiao, et al.
Published: (2024)
by: Wang, Zixiao, et al.
Published: (2024)
Customized Retrieval Augmented Generation and Benchmarking for EDA Tool Documentation QA
by: Pu, Yuan, et al.
Published: (2024)
by: Pu, Yuan, et al.
Published: (2024)
From Synthetic to Native: Benchmarking Multilingual Intent Classification in Logistics Customer Service
by: He, Haoyu, et al.
Published: (2026)
by: He, Haoyu, et al.
Published: (2026)
DV-World: Benchmarking Data Visualization Agents in Real-World Scenarios
by: Meng, Jinxiang, et al.
Published: (2026)
by: Meng, Jinxiang, et al.
Published: (2026)
DRIP-R: A Benchmark for Decision-Making and Reasoning Under Real-World Policy Ambiguity in the Retail Domain
by: Borkakoty, Hsuvas, et al.
Published: (2026)
by: Borkakoty, Hsuvas, et al.
Published: (2026)
Dial-In LLM: Human-Aligned LLM-in-the-loop Intent Clustering for Customer Service Dialogues
by: Hong, Mengze, et al.
Published: (2024)
by: Hong, Mengze, et al.
Published: (2024)
Can LLM Agents Simulate Multi-Turn Human Behavior? Evidence from Real Online Customer Behavior Data
by: Lu, Yuxuan, et al.
Published: (2025)
by: Lu, Yuxuan, et al.
Published: (2025)
AlpsBench: An LLM Personalization Benchmark for Real-Dialogue Memorization and Preference Alignment
by: Xiao, Jianfei, et al.
Published: (2026)
by: Xiao, Jianfei, et al.
Published: (2026)
Towards Proactive Personalization through Profile Customization for Individual Users in Dialogues
by: Zhang, Xiaotian, et al.
Published: (2025)
by: Zhang, Xiaotian, et al.
Published: (2025)
SEAD: Self-Evolving Agent for Multi-Turn Service Dialogue
by: Dai, Yuqin, et al.
Published: (2026)
by: Dai, Yuqin, et al.
Published: (2026)
Evaluating Bias in Spoken Dialogue LLMs for Real-World Decisions and Recommendations
by: Wu, Yihao, et al.
Published: (2025)
by: Wu, Yihao, et al.
Published: (2025)
RealMem: Benchmarking LLMs in Real-World Memory-Driven Interaction
by: Bian, Haonan, et al.
Published: (2026)
by: Bian, Haonan, et al.
Published: (2026)
Intent-driven In-context Learning for Few-shot Dialogue State Tracking
by: Yi, Zihao, et al.
Published: (2024)
by: Yi, Zihao, et al.
Published: (2024)
AERIC: Anticipatory Hidden-State Monitoring for Implicit Harmful Dialogue
by: Park, Jihyung, et al.
Published: (2026)
by: Park, Jihyung, et al.
Published: (2026)
PRSA: Prompt Stealing Attacks against Real-World Prompt Services
by: Yang, Yong, et al.
Published: (2024)
by: Yang, Yong, et al.
Published: (2024)
ECom-Bench: Can LLM Agent Resolve Real-World E-commerce Customer Support Issues?
by: Wang, Haoxin, et al.
Published: (2025)
by: Wang, Haoxin, et al.
Published: (2025)
RealFactBench: A Benchmark for Evaluating Large Language Models in Real-World Fact-Checking
by: Yang, Shuo, et al.
Published: (2025)
by: Yang, Shuo, et al.
Published: (2025)
RealTalk-CN: A Realistic Chinese Speech-Text Dialogue Benchmark With Cross-Modal Interaction Analysis
by: Wang, Enzhi, et al.
Published: (2025)
by: Wang, Enzhi, et al.
Published: (2025)
TravelPlanner: A Benchmark for Real-World Planning with Language Agents
by: Xie, Jian, et al.
Published: (2024)
by: Xie, Jian, et al.
Published: (2024)
StorySparkQA: Expert-Annotated QA Pairs with Real-World Knowledge for Children's Story-Based Learning
by: Chen, Jiaju, et al.
Published: (2023)
by: Chen, Jiaju, et al.
Published: (2023)
HalluDial: A Large-Scale Benchmark for Automatic Dialogue-Level Hallucination Evaluation
by: Luo, Wen, et al.
Published: (2024)
by: Luo, Wen, et al.
Published: (2024)
EvolMem: A Cognitive-Driven Benchmark for Multi-Session Dialogue Memory
by: Shen, Ye, et al.
Published: (2026)
by: Shen, Ye, et al.
Published: (2026)
ComperDial: Commonsense Persona-grounded Dialogue Dataset and Benchmark
by: Wakaki, Hiromi, et al.
Published: (2024)
by: Wakaki, Hiromi, et al.
Published: (2024)
RealBench: A Chinese Multi-image Understanding Benchmark Close to Real-world Scenarios
by: Zhao, Fei, et al.
Published: (2025)
by: Zhao, Fei, et al.
Published: (2025)
TableEval: A Real-World Benchmark for Complex, Multilingual, and Multi-Structured Table Question Answering
by: Zhu, Junnan, et al.
Published: (2025)
by: Zhu, Junnan, et al.
Published: (2025)
MARS-Bench: A Multi-turn Athletic Real-world Scenario Benchmark for Dialogue Evaluation
by: Yang, Chenghao, et al.
Published: (2025)
by: Yang, Chenghao, et al.
Published: (2025)
KnowMT-Bench: Benchmarking Knowledge-Intensive Long-Form Question Answering in Multi-Turn Dialogues
by: Chen, Junhao, et al.
Published: (2025)
by: Chen, Junhao, et al.
Published: (2025)
$τ$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
by: Yao, Shunyu, et al.
Published: (2024)
by: Yao, Shunyu, et al.
Published: (2024)
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue
by: Lù, Xing Han, et al.
Published: (2024)
by: Lù, Xing Han, et al.
Published: (2024)
WavBench: Benchmarking Reasoning, Colloquialism, and Paralinguistics for End-to-End Spoken Dialogue Models
by: Li, Yangzhuo, et al.
Published: (2026)
by: Li, Yangzhuo, et al.
Published: (2026)
Benchmarks Underestimate the Readiness of Multi-lingual Dialogue Agents
by: Lee, Andrew H., et al.
Published: (2024)
by: Lee, Andrew H., et al.
Published: (2024)
TaxPraBen: A Scalable Benchmark for Structured Evaluation of LLMs in Chinese Real-World Tax Practice
by: Hu, Gang, et al.
Published: (2026)
by: Hu, Gang, et al.
Published: (2026)
Planning, Creation, Usage: Benchmarking LLMs for Comprehensive Tool Utilization in Real-World Complex Scenarios
by: Huang, Shijue, et al.
Published: (2024)
by: Huang, Shijue, et al.
Published: (2024)
Underutilization of Syntactic Processing by Chinese Learners of English in Comprehending English Sentences, Evidenced from Adapted Garden-Path Ambiguity Experiment
by: Xu, Jiapeng
Published: (2024)
by: Xu, Jiapeng
Published: (2024)
MCP-SafetyBench: A Benchmark for Safety Evaluation of Large Language Models with Real-World MCP Servers
by: Zong, Xuanjun, et al.
Published: (2025)
by: Zong, Xuanjun, et al.
Published: (2025)
PLawBench: A Rubric-Based Benchmark for Evaluating LLMs in Real-World Legal Practice
by: Shi, Yuzhen, et al.
Published: (2026)
by: Shi, Yuzhen, et al.
Published: (2026)
Unstructured Text Enhanced Open-domain Dialogue System: A Systematic Survey
by: Ma, Longxuan, et al.
Published: (2024)
by: Ma, Longxuan, et al.
Published: (2024)
CEB: Compositional Evaluation Benchmark for Fairness in Large Language Models
by: Wang, Song, et al.
Published: (2024)
by: Wang, Song, et al.
Published: (2024)
CompliBench: Benchmarking LLM Judges for Compliance Violation Detection in Dialogue Systems
by: Yang, Jingbo, et al.
Published: (2026)
by: Yang, Jingbo, et al.
Published: (2026)
Similar Items
-
Reinforcing Real-world Service Agents: Balancing Utility and Cost in Task-oriented Dialogue
by: Gao, Ning, et al.
Published: (2026) -
ChatPattern: Layout Pattern Customization via Natural Language
by: Wang, Zixiao, et al.
Published: (2024) -
Customized Retrieval Augmented Generation and Benchmarking for EDA Tool Documentation QA
by: Pu, Yuan, et al.
Published: (2024) -
From Synthetic to Native: Benchmarking Multilingual Intent Classification in Logistics Customer Service
by: He, Haoyu, et al.
Published: (2026) -
DV-World: Benchmarking Data Visualization Agents in Real-World Scenarios
by: Meng, Jinxiang, et al.
Published: (2026)