The Behavior Gap: Evaluating Zero-shot LLM Agents in Complex Task-Oriented Dialogs
Fuente:
arXiv
Saved in:
| Main Authors: | Baidya, Avinash, Das, Kamalika, Gao, Xiang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Task-Oriented Dialog Systems for the Senegalese Wolof Language
by: Mbaye, Derguene, et al.
Published: (2024)
by: Mbaye, Derguene, et al.
Published: (2024)
Agent Laboratory: Using LLM Agents as Research Assistants
by: Schmidgall, Samuel, et al.
Published: (2025)
by: Schmidgall, Samuel, et al.
Published: (2025)
Policy Maps: Tools for Guiding the Unbounded Space of LLM Behaviors
by: Lam, Michelle S., et al.
Published: (2024)
by: Lam, Michelle S., et al.
Published: (2024)
Multimodal Behavioral Patterns Analysis with Eye-Tracking and LLM-Based Reasoning
by: Guo, Dongyang, et al.
Published: (2025)
by: Guo, Dongyang, et al.
Published: (2025)
Cross-Lingual Prompt Steerability: Towards Accurate and Robust LLM Behavior across Languages
by: Zhang, Lechen, et al.
Published: (2025)
by: Zhang, Lechen, et al.
Published: (2025)
Detecting and Preventing Harmful Behaviors in AI Companions: Development and Evaluation of the SHIELD Supervisory System
by: Ben-Zion, Ziv, et al.
Published: (2025)
by: Ben-Zion, Ziv, et al.
Published: (2025)
How Controllable Are Large Language Models? A Unified Evaluation across Behavioral Granularities
by: Xu, Ziwen, et al.
Published: (2026)
by: Xu, Ziwen, et al.
Published: (2026)
Evaluation of LLMs-based Hidden States as Author Representations for Psychological Human-Centered NLP Tasks
by: Soni, Nikita, et al.
Published: (2025)
by: Soni, Nikita, et al.
Published: (2025)
TICKing All the Boxes: Generated Checklists Improve LLM Evaluation and Generation
by: Cook, Jonathan, et al.
Published: (2024)
by: Cook, Jonathan, et al.
Published: (2024)
DigiData: Training and Evaluating General-Purpose Mobile Control Agents
by: Sun, Yuxuan, et al.
Published: (2025)
by: Sun, Yuxuan, et al.
Published: (2025)
VAL: Interactive Task Learning with GPT Dialog Parsing
by: Lawley, Lane, et al.
Published: (2023)
by: Lawley, Lane, et al.
Published: (2023)
LLM Comparator: Visual Analytics for Side-by-Side Evaluation of Large Language Models
by: Kahng, Minsuk, et al.
Published: (2024)
by: Kahng, Minsuk, et al.
Published: (2024)
Can Generative AI Support Patients' & Caregivers' Informational Needs? Towards Task-Centric Evaluation Of AI Systems
by: Rajagopal, Shreya, et al.
Published: (2024)
by: Rajagopal, Shreya, et al.
Published: (2024)
Building Trust in Mental Health Chatbots: Safety Metrics and LLM-Based Evaluation Tools
by: Park, Jung In, et al.
Published: (2024)
by: Park, Jung In, et al.
Published: (2024)
Augmenting Human Evaluation with LLM Judges: How Many Human Reviews Do You Need?
by: Kim, Jane Paik
Published: (2026)
by: Kim, Jane Paik
Published: (2026)
Zero-shot causal learning
by: Nilforoshan, Hamed, et al.
Published: (2023)
by: Nilforoshan, Hamed, et al.
Published: (2023)
Never Start from Scratch: Expediting On-Device LLM Personalization via Explainable Model Selection
by: Wang, Haoming, et al.
Published: (2025)
by: Wang, Haoming, et al.
Published: (2025)
The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas
by: Si, Chenglei, et al.
Published: (2025)
by: Si, Chenglei, et al.
Published: (2025)
Conversation Routines: A Prompt Engineering Framework for Task-Oriented Dialog Systems
by: Robino, Giorgio
Published: (2025)
by: Robino, Giorgio
Published: (2025)
Collaborative Causal Sensemaking: Closing the Complementarity Gap in Human-AI Decision Support
by: Jain, Raunak
Published: (2025)
by: Jain, Raunak
Published: (2025)
Feedback-Aware Monte Carlo Tree Search for Efficient Information Seeking in Goal-Oriented Conversations
by: Chopra, Harshita, et al.
Published: (2025)
by: Chopra, Harshita, et al.
Published: (2025)
Vibe Check: Understanding the Effects of LLM-Based Conversational Agents' Personality and Alignment on User Perceptions in Goal-Oriented Tasks
by: Rahman, Hasibur, et al.
Published: (2025)
by: Rahman, Hasibur, et al.
Published: (2025)
LLM Attributor: Interactive Visual Attribution for LLM Generation
by: Lee, Seongmin, et al.
Published: (2024)
by: Lee, Seongmin, et al.
Published: (2024)
Are You Being Tracked? Discover the Power of Zero-Shot Trajectory Tracing with LLMs!
by: Yang, Huanqi, et al.
Published: (2024)
by: Yang, Huanqi, et al.
Published: (2024)
AutoGLM: Autonomous Foundation Agents for GUIs
by: Liu, Xiao, et al.
Published: (2024)
by: Liu, Xiao, et al.
Published: (2024)
Observing Dialogue in Therapy: Categorizing and Forecasting Behavioral Codes
by: Cao, Jie, et al.
Published: (2019)
by: Cao, Jie, et al.
Published: (2019)
KnowAgent: Knowledge-Augmented Planning for LLM-Based Agents
by: Zhu, Yuqi, et al.
Published: (2024)
by: Zhu, Yuqi, et al.
Published: (2024)
Evaluating Behavioral Alignment in Conflict Dialogue: A Multi-Dimensional Comparison of LLM Agents and Humans
by: Kwon, Deuksin, et al.
Published: (2025)
by: Kwon, Deuksin, et al.
Published: (2025)
DiscoverLLM: From Executing Intents to Discovering Them
by: Kim, Tae Soo, et al.
Published: (2026)
by: Kim, Tae Soo, et al.
Published: (2026)
MotionTeller: Multi-modal Integration of Wearable Time-Series with LLMs for Health and Behavioral Understanding
by: Zhang, Aiwei, et al.
Published: (2025)
by: Zhang, Aiwei, et al.
Published: (2025)
Estimating LLM Consistency: A User Baseline vs Surrogate Metrics
by: Wu, Xiaoyuan, et al.
Published: (2025)
by: Wu, Xiaoyuan, et al.
Published: (2025)
Creative Beam Search: LLM-as-a-Judge For Improving Response Generation
by: Franceschelli, Giorgio, et al.
Published: (2024)
by: Franceschelli, Giorgio, et al.
Published: (2024)
Language Models as Zero-Shot Trajectory Generators
by: Kwon, Teyun, et al.
Published: (2023)
by: Kwon, Teyun, et al.
Published: (2023)
Talking with Oompa Loompas: A novel framework for evaluating linguistic acquisition of LLM agents
by: Swain, Sankalp Tattwadarshi, et al.
Published: (2025)
by: Swain, Sankalp Tattwadarshi, et al.
Published: (2025)
Heterogeneous Value Alignment Evaluation for Large Language Models
by: Zhang, Zhaowei, et al.
Published: (2023)
by: Zhang, Zhaowei, et al.
Published: (2023)
Improving Dialogue Agents by Decomposing One Global Explicit Annotation with Local Implicit Multimodal Feedback
by: Lee, Dong Won, et al.
Published: (2024)
by: Lee, Dong Won, et al.
Published: (2024)
PREF: Reference-Free Evaluation of Personalised Text Generation in LLMs
by: Fu, Xiao, et al.
Published: (2025)
by: Fu, Xiao, et al.
Published: (2025)
Evaluating Large Language Models for Health-related Queries with Presuppositions
by: Kaur, Navreet, et al.
Published: (2023)
by: Kaur, Navreet, et al.
Published: (2023)
Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025
by: Thakkar, Nitya, et al.
Published: (2025)
by: Thakkar, Nitya, et al.
Published: (2025)
AIRepr: An Analyst-Inspector Framework for Evaluating Reproducibility of LLMs in Data Science
by: Zeng, Qiuhai, et al.
Published: (2025)
by: Zeng, Qiuhai, et al.
Published: (2025)
Similar Items
-
Task-Oriented Dialog Systems for the Senegalese Wolof Language
by: Mbaye, Derguene, et al.
Published: (2024) -
Agent Laboratory: Using LLM Agents as Research Assistants
by: Schmidgall, Samuel, et al.
Published: (2025) -
Policy Maps: Tools for Guiding the Unbounded Space of LLM Behaviors
by: Lam, Michelle S., et al.
Published: (2024) -
Multimodal Behavioral Patterns Analysis with Eye-Tracking and LLM-Based Reasoning
by: Guo, Dongyang, et al.
Published: (2025) -
Cross-Lingual Prompt Steerability: Towards Accurate and Robust LLM Behavior across Languages
by: Zhang, Lechen, et al.
Published: (2025)