Measuring and Mitigating the Distributional Gap Between Real and Simulated User Behaviors
Fuente:
arXiv
Saved in:
| Main Authors: | Mehri, Shuhaib, Laban, Philippe, Shashidhar, Sumuk, Abdulhai, Marwa, Levine, Sergey, Galley, Michel, Hakkani-Tür, Dilek |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Question Generation for Assessing Early Literacy Reading Comprehension
by: Yang, Xiaocheng, et al.
Published: (2025)
by: Yang, Xiaocheng, et al.
Published: (2025)
Goal Alignment in LLM-Based User Simulators for Conversational AI
by: Mehri, Shuhaib, et al.
Published: (2025)
by: Mehri, Shuhaib, et al.
Published: (2025)
AURA: A Diagnostic Framework for Tracking User Satisfaction of Interactive Planning Agents
by: Kim, Takyoung, et al.
Published: (2025)
by: Kim, Takyoung, et al.
Published: (2025)
Unsupervised Human Preference Learning
by: Shashidhar, Sumuk, et al.
Published: (2024)
by: Shashidhar, Sumuk, et al.
Published: (2024)
Beyond Sample-Level Feedback: Using Reference-Level Feedback to Guide Data Synthesis
by: Mehri, Shuhaib, et al.
Published: (2025)
by: Mehri, Shuhaib, et al.
Published: (2025)
Sparking Scientific Creativity via LLM-Driven Interdisciplinary Inspiration
by: Kargupta, Priyanka, et al.
Published: (2026)
by: Kargupta, Priyanka, et al.
Published: (2026)
Persuade Me if You Can: A Framework for Evaluating Persuasion Effectiveness and Susceptibility Among Large Language Models
by: Bozdag, Nimet Beyza, et al.
Published: (2025)
by: Bozdag, Nimet Beyza, et al.
Published: (2025)
User Preference Modeling for Conversational LLM Agents: Weak Rewards from Retrieval-Augmented Interaction
by: Hao, Yuren, et al.
Published: (2026)
by: Hao, Yuren, et al.
Published: (2026)
YourBench: Easy Custom Evaluation Sets for Everyone
by: Shashidhar, Sumuk, et al.
Published: (2025)
by: Shashidhar, Sumuk, et al.
Published: (2025)
MultiSessionCollab: Learning User Preferences with Memory to Improve Long-Term Collaboration
by: Mehri, Shuhaib, et al.
Published: (2026)
by: Mehri, Shuhaib, et al.
Published: (2026)
Simulating User Agents for Embodied Conversational-AI
by: Philipov, Daniel, et al.
Published: (2024)
by: Philipov, Daniel, et al.
Published: (2024)
PSI-Bench: Towards Clinically Grounded and Interpretable Evaluation of Depression Patient Simulators
by: Hoang, Nguyen Khoi, et al.
Published: (2026)
by: Hoang, Nguyen Khoi, et al.
Published: (2026)
Do LLMs Encode Functional Importance of Reasoning Tokens?
by: Singh, Janvijay, et al.
Published: (2026)
by: Singh, Janvijay, et al.
Published: (2026)
Neural Networks for Learnable and Scalable Influence Estimation of Instruction Fine-Tuning Data
by: Agarwal, Ishika, et al.
Published: (2025)
by: Agarwal, Ishika, et al.
Published: (2025)
Embodied Multi-Agent Coordination by Aligning World Models Through Dialogue
by: Dongre, Vardhan, et al.
Published: (2026)
by: Dongre, Vardhan, et al.
Published: (2026)
Large Language Models as User-Agents for Evaluating Task-Oriented-Dialogue Systems
by: Kazi, Taaha, et al.
Published: (2024)
by: Kazi, Taaha, et al.
Published: (2024)
Must Read: A Comprehensive Survey of Computational Persuasion
by: Bozdag, Nimet Beyza, et al.
Published: (2025)
by: Bozdag, Nimet Beyza, et al.
Published: (2025)
Spark: A System for Scientifically Creative Idea Generation
by: Sanyal, Aishik, et al.
Published: (2025)
by: Sanyal, Aishik, et al.
Published: (2025)
Consistently Simulating Human Personas with Multi-Turn Reinforcement Learning
by: Abdulhai, Marwa, et al.
Published: (2025)
by: Abdulhai, Marwa, et al.
Published: (2025)
Know Your Mistakes: Towards Preventing Overreliance on Task-Oriented Conversational AI Through Accountability Modeling
by: Dey, Suvodip, et al.
Published: (2025)
by: Dey, Suvodip, et al.
Published: (2025)
Confidence Estimation for LLM-Based Dialogue State Tracking
by: Sun, Yi-Jyun, et al.
Published: (2024)
by: Sun, Yi-Jyun, et al.
Published: (2024)
From Fact to Judgment: Investigating the Impact of Task Framing on LLM Conviction in Dialogue Systems
by: Rabbani, Parisa, et al.
Published: (2025)
by: Rabbani, Parisa, et al.
Published: (2025)
Combinatorial Creativity: A New Frontier in Generalization Abilities
by: Schapiro, Samuel, et al.
Published: (2025)
by: Schapiro, Samuel, et al.
Published: (2025)
A Rising Tide Lifts All Boats: MTQE Rewards for Idioms Improve General Translation Quality
by: Agarwal, Ishika, et al.
Published: (2026)
by: Agarwal, Ishika, et al.
Published: (2026)
LLMs are Vulnerable to Malicious Prompts Disguised as Scientific Language
by: Ge, Yubin, et al.
Published: (2025)
by: Ge, Yubin, et al.
Published: (2025)
ReasoningFlow: Semantic Structure of Complex Reasoning Traces
by: Lee, Jinu, et al.
Published: (2025)
by: Lee, Jinu, et al.
Published: (2025)
Instruct, Not Assist: LLM-based Multi-Turn Planning and Hierarchical Questioning for Socratic Code Debugging
by: Kargupta, Priyanka, et al.
Published: (2024)
by: Kargupta, Priyanka, et al.
Published: (2024)
SMART: Self-Aware Agent for Tool Overuse Mitigation
by: Qian, Cheng, et al.
Published: (2025)
by: Qian, Cheng, et al.
Published: (2025)
Language Specific Knowledge: Do Models Know Better in X than in English?
by: Agarwal, Ishika, et al.
Published: (2025)
by: Agarwal, Ishika, et al.
Published: (2025)
Self-Improving LLM Agents at Test-Time
by: Acikgoz, Emre Can, et al.
Published: (2025)
by: Acikgoz, Emre Can, et al.
Published: (2025)
Evaluating & Reducing Deceptive Dialogue From Language Models with Multi-turn RL
by: Abdulhai, Marwa, et al.
Published: (2025)
by: Abdulhai, Marwa, et al.
Published: (2025)
DialDefer: A Framework for Detecting and Mitigating LLM Dialogic Deference
by: Rabbani, Parisa, et al.
Published: (2026)
by: Rabbani, Parisa, et al.
Published: (2026)
Premise-Augmented Reasoning Chains Improve Error Identification in Math reasoning with LLMs
by: Mukherjee, Sagnik, et al.
Published: (2025)
by: Mukherjee, Sagnik, et al.
Published: (2025)
DocCHA: Towards LLM-Augmented Interactive Online diagnosis System
by: Liu, Xinyi, et al.
Published: (2025)
by: Liu, Xinyi, et al.
Published: (2025)
ReSpAct: Harmonizing Reasoning, Speaking, and Acting Towards Building Large Language Model-Based Conversational AI Agents
by: Dongre, Vardhan, et al.
Published: (2024)
by: Dongre, Vardhan, et al.
Published: (2024)
ATOD: An Evaluation Framework and Benchmark for Agentic Task-Oriented Dialogue Systems
by: Zhang, Yifei, et al.
Published: (2026)
by: Zhang, Yifei, et al.
Published: (2026)
TD-EVAL: Revisiting Task-Oriented Dialogue Evaluation by Combining Turn-Level Precision with Dialogue-Level Comparisons
by: Acikgoz, Emre Can, et al.
Published: (2025)
by: Acikgoz, Emre Can, et al.
Published: (2025)
Enabling Chatbots with Eyes and Ears: An Immersive Multimodal Conversation System for Dynamic Interactions
by: Jang, Jihyoung, et al.
Published: (2025)
by: Jang, Jihyoung, et al.
Published: (2025)
MAC: A Multi-Agent Framework for Interactive User Clarification in Multi-turn Conversations
by: Acikgoz, Emre Can, et al.
Published: (2025)
by: Acikgoz, Emre Can, et al.
Published: (2025)
Dialog Flow Induction for Constrainable LLM-Based Chatbots
by: Agrawal, Stuti, et al.
Published: (2024)
by: Agrawal, Stuti, et al.
Published: (2024)
Similar Items
-
Question Generation for Assessing Early Literacy Reading Comprehension
by: Yang, Xiaocheng, et al.
Published: (2025) -
Goal Alignment in LLM-Based User Simulators for Conversational AI
by: Mehri, Shuhaib, et al.
Published: (2025) -
AURA: A Diagnostic Framework for Tracking User Satisfaction of Interactive Planning Agents
by: Kim, Takyoung, et al.
Published: (2025) -
Unsupervised Human Preference Learning
by: Shashidhar, Sumuk, et al.
Published: (2024) -
Beyond Sample-Level Feedback: Using Reference-Level Feedback to Guide Data Synthesis
by: Mehri, Shuhaib, et al.
Published: (2025)