Beyond IVR: Benchmarking Customer Support LLM Agents for Business-Adherence
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Balaji, Sumanth, Mishra, Piyush, Sachdeva, Aashraya, Agrawal, Suraj |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
RECOVER: Robust Entity Correction via agentic Orchestration of hypothesis Variants for Evidence-based Recovery
par: Kumar, Abhishek, et autres
Publié: (2026)
par: Kumar, Abhishek, et autres
Publié: (2026)
Beyond IVR Touch-Tones: Customer Intent Routing using LLMs
par: Rojas-Galeano, Sergio
Publié: (2025)
par: Rojas-Galeano, Sergio
Publié: (2025)
Beyond Bias Scores: Unmasking Vacuous Neutrality in Small Language Models
par: Manduru, Sumanth, et autres
Publié: (2025)
par: Manduru, Sumanth, et autres
Publié: (2025)
BeliefShift: Benchmarking Temporal Belief Consistency and Opinion Drift in LLM Agents
par: Myakala, Praveen Kumar, et autres
Publié: (2026)
par: Myakala, Praveen Kumar, et autres
Publié: (2026)
Beyond the Rubric: Cultural Misalignment in LLM Benchmarks for Sexual and Reproductive Health
par: Dey, Sumon Kanti, et autres
Publié: (2025)
par: Dey, Sumon Kanti, et autres
Publié: (2025)
MindFlow: Revolutionizing E-commerce Customer Support with Multimodal LLM Agents
par: Gong, Ming, et autres
Publié: (2025)
par: Gong, Ming, et autres
Publié: (2025)
Beyond Sentiment: A Multi-Agent Pipeline for Actionable Business Advice from Reviews
par: Bhandari, Kartikey Singh, et autres
Publié: (2026)
par: Bhandari, Kartikey Singh, et autres
Publié: (2026)
ECom-Bench: Can LLM Agent Resolve Real-World E-commerce Customer Support Issues?
par: Wang, Haoxin, et autres
Publié: (2025)
par: Wang, Haoxin, et autres
Publié: (2025)
PEDAL: Enhancing Greedy Decoding with Large Language Models using Diverse Exemplars
par: Prabhu, Sumanth
Publié: (2024)
par: Prabhu, Sumanth
Publié: (2024)
A Benchmark Dataset and Evaluation Framework for Vietnamese Large Language Models in Customer Support
par: Nguyen, Long S. T., et autres
Publié: (2025)
par: Nguyen, Long S. T., et autres
Publié: (2025)
Beyond Perplexity: Character Distribution Signatures and the MDTA Benchmark for AI Text Detection
par: Narayanasamy, Priyadarshan, et autres
Publié: (2026)
par: Narayanasamy, Priyadarshan, et autres
Publié: (2026)
JudgeAgent: Beyond Static Benchmarks for Knowledge-Driven and Dynamic LLM Evaluation
par: Shi, Zhichao, et autres
Publié: (2025)
par: Shi, Zhichao, et autres
Publié: (2025)
Beyond Next Word Prediction: Developing Comprehensive Evaluation Frameworks for measuring LLM performance on real world applications
par: Agrawal, Vishakha, et autres
Publié: (2025)
par: Agrawal, Vishakha, et autres
Publié: (2025)
GFlowVLM: Enhancing Multi-step Reasoning in Vision-Language Models with Generative Flow Networks
par: Kang, Haoqiang, et autres
Publié: (2025)
par: Kang, Haoqiang, et autres
Publié: (2025)
Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory
par: Wei, Tianxin, et autres
Publié: (2025)
par: Wei, Tianxin, et autres
Publié: (2025)
Improving LLM Safety and Helpfulness using SFT and DPO: A Study on OPT-350M
par: Pant, Piyush
Publié: (2025)
par: Pant, Piyush
Publié: (2025)
Evaluating, Synthesizing, and Enhancing for Customer Support Conversation
par: Zhu, Jie, et autres
Publié: (2025)
par: Zhu, Jie, et autres
Publié: (2025)
Accelerating Direct Preference Optimization with Prefix Sharing
par: Wang, Franklin, et autres
Publié: (2024)
par: Wang, Franklin, et autres
Publié: (2024)
Agent Bain vs. Agent McKinsey: A New Text-to-SQL Benchmark for the Business Domain
par: Li, Yue, et autres
Publié: (2025)
par: Li, Yue, et autres
Publié: (2025)
LLM Agent Meets Agentic AI: Can LLM Agents Simulate Customers to Evaluate Agentic-AI-based Shopping Assistants?
par: Sun, Lu, et autres
Publié: (2025)
par: Sun, Lu, et autres
Publié: (2025)
Learning Selective LLM Autonomy from Copilot Feedback in Enterprise Customer Support Workflows
par: Borovkov, Nikita, et autres
Publié: (2026)
par: Borovkov, Nikita, et autres
Publié: (2026)
AgentHallu: Benchmarking Automated Hallucination Attribution of LLM-based Agents
par: Liu, Xuannan, et autres
Publié: (2026)
par: Liu, Xuannan, et autres
Publié: (2026)
ENPMR-Bench: Benchmarking Proactive Memory Retrieval for Emotional Support Agents
par: Fu, Xing, et autres
Publié: (2026)
par: Fu, Xing, et autres
Publié: (2026)
Incremental Summarization for Customer Support via Progressive Note-Taking and Agent Feedback
par: Wu, Yisha, et autres
Publié: (2025)
par: Wu, Yisha, et autres
Publié: (2025)
Pluralistic Behavior Suite: Stress-Testing Multi-Turn Adherence to Custom Behavioral Policies
par: Varshney, Prasoon, et autres
Publié: (2025)
par: Varshney, Prasoon, et autres
Publié: (2025)
Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks
par: Cao, Yixin, et autres
Publié: (2025)
par: Cao, Yixin, et autres
Publié: (2025)
TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks
par: Xu, Frank F., et autres
Publié: (2024)
par: Xu, Frank F., et autres
Publié: (2024)
When Agents Trade: Live Multi-Market Trading Benchmark for LLM Agents
par: Qian, Lingfei, et autres
Publié: (2025)
par: Qian, Lingfei, et autres
Publié: (2025)
LLMRank: Understanding LLM Strengths for Model Routing
par: Agrawal, Shubham, et autres
Publié: (2025)
par: Agrawal, Shubham, et autres
Publié: (2025)
Benchmarking and Learning Real-World Customer Service Dialogue
par: Gao, Tianhong, et autres
Publié: (2025)
par: Gao, Tianhong, et autres
Publié: (2025)
Sustainable Digitalization of Business with Multi-Agent RAG and LLM
par: Arslan, Muhammad, et autres
Publié: (2025)
par: Arslan, Muhammad, et autres
Publié: (2025)
Customer-R1: Personalized Simulation of Human Behaviors via RL-based LLM Agent in Online Shopping
par: Wang, Ziyi, et autres
Publié: (2025)
par: Wang, Ziyi, et autres
Publié: (2025)
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents
par: Shen, Yujiong, et autres
Publié: (2026)
par: Shen, Yujiong, et autres
Publié: (2026)
MALIBU Benchmark: Multi-Agent LLM Implicit Bias Uncovered
par: Mirza, Imran, et autres
Publié: (2025)
par: Mirza, Imran, et autres
Publié: (2025)
Business as Rulesual: A Benchmark and Framework for Business Rule Flow Modeling with LLMs
par: Yang, Chen, et autres
Publié: (2025)
par: Yang, Chen, et autres
Publié: (2025)
GUIDEQ: Framework for Guided Questioning for progressive informational collection and classification
par: Mishra, Priya, et autres
Publié: (2024)
par: Mishra, Priya, et autres
Publié: (2024)
CharacterBench: Benchmarking Character Customization of Large Language Models
par: Zhou, Jinfeng, et autres
Publié: (2024)
par: Zhou, Jinfeng, et autres
Publié: (2024)
Vocabulary Customization for Efficient Domain-Specific LLM Deployment
par: Herold, Christian, et autres
Publié: (2025)
par: Herold, Christian, et autres
Publié: (2025)
Can LLM Agents Simulate Multi-Turn Human Behavior? Evidence from Real Online Customer Behavior Data
par: Lu, Yuxuan, et autres
Publié: (2025)
par: Lu, Yuxuan, et autres
Publié: (2025)
Building a Few-Shot Cross-Domain Multilingual NLU Model for Customer Care
par: Kumar, Saurabh, et autres
Publié: (2025)
par: Kumar, Saurabh, et autres
Publié: (2025)
Documents similaires
-
RECOVER: Robust Entity Correction via agentic Orchestration of hypothesis Variants for Evidence-based Recovery
par: Kumar, Abhishek, et autres
Publié: (2026) -
Beyond IVR Touch-Tones: Customer Intent Routing using LLMs
par: Rojas-Galeano, Sergio
Publié: (2025) -
Beyond Bias Scores: Unmasking Vacuous Neutrality in Small Language Models
par: Manduru, Sumanth, et autres
Publié: (2025) -
BeliefShift: Benchmarking Temporal Belief Consistency and Opinion Drift in LLM Agents
par: Myakala, Praveen Kumar, et autres
Publié: (2026) -
Beyond the Rubric: Cultural Misalignment in LLM Benchmarks for Sexual and Reproductive Health
par: Dey, Sumon Kanti, et autres
Publié: (2025)