Benchmarking and Learning Real-World Customer Service Dialogue
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gao, Tianhong, Shen, Jundong, Wang, Jiapeng, Shi, Bei, Ju, Ying, Yao, Junfeng, Yu, Huiyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Reinforcing Real-world Service Agents: Balancing Utility and Cost in Task-oriented Dialogue
von: Gao, Ning, et al.
Veröffentlicht: (2026)
von: Gao, Ning, et al.
Veröffentlicht: (2026)
ChatPattern: Layout Pattern Customization via Natural Language
von: Wang, Zixiao, et al.
Veröffentlicht: (2024)
von: Wang, Zixiao, et al.
Veröffentlicht: (2024)
Customized Retrieval Augmented Generation and Benchmarking for EDA Tool Documentation QA
von: Pu, Yuan, et al.
Veröffentlicht: (2024)
von: Pu, Yuan, et al.
Veröffentlicht: (2024)
From Synthetic to Native: Benchmarking Multilingual Intent Classification in Logistics Customer Service
von: He, Haoyu, et al.
Veröffentlicht: (2026)
von: He, Haoyu, et al.
Veröffentlicht: (2026)
DV-World: Benchmarking Data Visualization Agents in Real-World Scenarios
von: Meng, Jinxiang, et al.
Veröffentlicht: (2026)
von: Meng, Jinxiang, et al.
Veröffentlicht: (2026)
DRIP-R: A Benchmark for Decision-Making and Reasoning Under Real-World Policy Ambiguity in the Retail Domain
von: Borkakoty, Hsuvas, et al.
Veröffentlicht: (2026)
von: Borkakoty, Hsuvas, et al.
Veröffentlicht: (2026)
Dial-In LLM: Human-Aligned LLM-in-the-loop Intent Clustering for Customer Service Dialogues
von: Hong, Mengze, et al.
Veröffentlicht: (2024)
von: Hong, Mengze, et al.
Veröffentlicht: (2024)
Can LLM Agents Simulate Multi-Turn Human Behavior? Evidence from Real Online Customer Behavior Data
von: Lu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Lu, Yuxuan, et al.
Veröffentlicht: (2025)
AlpsBench: An LLM Personalization Benchmark for Real-Dialogue Memorization and Preference Alignment
von: Xiao, Jianfei, et al.
Veröffentlicht: (2026)
von: Xiao, Jianfei, et al.
Veröffentlicht: (2026)
Towards Proactive Personalization through Profile Customization for Individual Users in Dialogues
von: Zhang, Xiaotian, et al.
Veröffentlicht: (2025)
von: Zhang, Xiaotian, et al.
Veröffentlicht: (2025)
SEAD: Self-Evolving Agent for Multi-Turn Service Dialogue
von: Dai, Yuqin, et al.
Veröffentlicht: (2026)
von: Dai, Yuqin, et al.
Veröffentlicht: (2026)
Evaluating Bias in Spoken Dialogue LLMs for Real-World Decisions and Recommendations
von: Wu, Yihao, et al.
Veröffentlicht: (2025)
von: Wu, Yihao, et al.
Veröffentlicht: (2025)
RealMem: Benchmarking LLMs in Real-World Memory-Driven Interaction
von: Bian, Haonan, et al.
Veröffentlicht: (2026)
von: Bian, Haonan, et al.
Veröffentlicht: (2026)
Intent-driven In-context Learning for Few-shot Dialogue State Tracking
von: Yi, Zihao, et al.
Veröffentlicht: (2024)
von: Yi, Zihao, et al.
Veröffentlicht: (2024)
AERIC: Anticipatory Hidden-State Monitoring for Implicit Harmful Dialogue
von: Park, Jihyung, et al.
Veröffentlicht: (2026)
von: Park, Jihyung, et al.
Veröffentlicht: (2026)
PRSA: Prompt Stealing Attacks against Real-World Prompt Services
von: Yang, Yong, et al.
Veröffentlicht: (2024)
von: Yang, Yong, et al.
Veröffentlicht: (2024)
ECom-Bench: Can LLM Agent Resolve Real-World E-commerce Customer Support Issues?
von: Wang, Haoxin, et al.
Veröffentlicht: (2025)
von: Wang, Haoxin, et al.
Veröffentlicht: (2025)
RealFactBench: A Benchmark for Evaluating Large Language Models in Real-World Fact-Checking
von: Yang, Shuo, et al.
Veröffentlicht: (2025)
von: Yang, Shuo, et al.
Veröffentlicht: (2025)
RealTalk-CN: A Realistic Chinese Speech-Text Dialogue Benchmark With Cross-Modal Interaction Analysis
von: Wang, Enzhi, et al.
Veröffentlicht: (2025)
von: Wang, Enzhi, et al.
Veröffentlicht: (2025)
TravelPlanner: A Benchmark for Real-World Planning with Language Agents
von: Xie, Jian, et al.
Veröffentlicht: (2024)
von: Xie, Jian, et al.
Veröffentlicht: (2024)
StorySparkQA: Expert-Annotated QA Pairs with Real-World Knowledge for Children's Story-Based Learning
von: Chen, Jiaju, et al.
Veröffentlicht: (2023)
von: Chen, Jiaju, et al.
Veröffentlicht: (2023)
HalluDial: A Large-Scale Benchmark for Automatic Dialogue-Level Hallucination Evaluation
von: Luo, Wen, et al.
Veröffentlicht: (2024)
von: Luo, Wen, et al.
Veröffentlicht: (2024)
EvolMem: A Cognitive-Driven Benchmark for Multi-Session Dialogue Memory
von: Shen, Ye, et al.
Veröffentlicht: (2026)
von: Shen, Ye, et al.
Veröffentlicht: (2026)
ComperDial: Commonsense Persona-grounded Dialogue Dataset and Benchmark
von: Wakaki, Hiromi, et al.
Veröffentlicht: (2024)
von: Wakaki, Hiromi, et al.
Veröffentlicht: (2024)
RealBench: A Chinese Multi-image Understanding Benchmark Close to Real-world Scenarios
von: Zhao, Fei, et al.
Veröffentlicht: (2025)
von: Zhao, Fei, et al.
Veröffentlicht: (2025)
TableEval: A Real-World Benchmark for Complex, Multilingual, and Multi-Structured Table Question Answering
von: Zhu, Junnan, et al.
Veröffentlicht: (2025)
von: Zhu, Junnan, et al.
Veröffentlicht: (2025)
MARS-Bench: A Multi-turn Athletic Real-world Scenario Benchmark for Dialogue Evaluation
von: Yang, Chenghao, et al.
Veröffentlicht: (2025)
von: Yang, Chenghao, et al.
Veröffentlicht: (2025)
KnowMT-Bench: Benchmarking Knowledge-Intensive Long-Form Question Answering in Multi-Turn Dialogues
von: Chen, Junhao, et al.
Veröffentlicht: (2025)
von: Chen, Junhao, et al.
Veröffentlicht: (2025)
$τ$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
von: Yao, Shunyu, et al.
Veröffentlicht: (2024)
von: Yao, Shunyu, et al.
Veröffentlicht: (2024)
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue
von: Lù, Xing Han, et al.
Veröffentlicht: (2024)
von: Lù, Xing Han, et al.
Veröffentlicht: (2024)
WavBench: Benchmarking Reasoning, Colloquialism, and Paralinguistics for End-to-End Spoken Dialogue Models
von: Li, Yangzhuo, et al.
Veröffentlicht: (2026)
von: Li, Yangzhuo, et al.
Veröffentlicht: (2026)
Benchmarks Underestimate the Readiness of Multi-lingual Dialogue Agents
von: Lee, Andrew H., et al.
Veröffentlicht: (2024)
von: Lee, Andrew H., et al.
Veröffentlicht: (2024)
TaxPraBen: A Scalable Benchmark for Structured Evaluation of LLMs in Chinese Real-World Tax Practice
von: Hu, Gang, et al.
Veröffentlicht: (2026)
von: Hu, Gang, et al.
Veröffentlicht: (2026)
Planning, Creation, Usage: Benchmarking LLMs for Comprehensive Tool Utilization in Real-World Complex Scenarios
von: Huang, Shijue, et al.
Veröffentlicht: (2024)
von: Huang, Shijue, et al.
Veröffentlicht: (2024)
Underutilization of Syntactic Processing by Chinese Learners of English in Comprehending English Sentences, Evidenced from Adapted Garden-Path Ambiguity Experiment
von: Xu, Jiapeng
Veröffentlicht: (2024)
von: Xu, Jiapeng
Veröffentlicht: (2024)
MCP-SafetyBench: A Benchmark for Safety Evaluation of Large Language Models with Real-World MCP Servers
von: Zong, Xuanjun, et al.
Veröffentlicht: (2025)
von: Zong, Xuanjun, et al.
Veröffentlicht: (2025)
PLawBench: A Rubric-Based Benchmark for Evaluating LLMs in Real-World Legal Practice
von: Shi, Yuzhen, et al.
Veröffentlicht: (2026)
von: Shi, Yuzhen, et al.
Veröffentlicht: (2026)
Unstructured Text Enhanced Open-domain Dialogue System: A Systematic Survey
von: Ma, Longxuan, et al.
Veröffentlicht: (2024)
von: Ma, Longxuan, et al.
Veröffentlicht: (2024)
CEB: Compositional Evaluation Benchmark for Fairness in Large Language Models
von: Wang, Song, et al.
Veröffentlicht: (2024)
von: Wang, Song, et al.
Veröffentlicht: (2024)
CompliBench: Benchmarking LLM Judges for Compliance Violation Detection in Dialogue Systems
von: Yang, Jingbo, et al.
Veröffentlicht: (2026)
von: Yang, Jingbo, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Reinforcing Real-world Service Agents: Balancing Utility and Cost in Task-oriented Dialogue
von: Gao, Ning, et al.
Veröffentlicht: (2026) -
ChatPattern: Layout Pattern Customization via Natural Language
von: Wang, Zixiao, et al.
Veröffentlicht: (2024) -
Customized Retrieval Augmented Generation and Benchmarking for EDA Tool Documentation QA
von: Pu, Yuan, et al.
Veröffentlicht: (2024) -
From Synthetic to Native: Benchmarking Multilingual Intent Classification in Logistics Customer Service
von: He, Haoyu, et al.
Veröffentlicht: (2026) -
DV-World: Benchmarking Data Visualization Agents in Real-World Scenarios
von: Meng, Jinxiang, et al.
Veröffentlicht: (2026)