SimulatorArena: Are User Simulators Reliable Proxies for Multi-Turn Evaluation of AI Assistants?
Fuente:
arXiv
Saved in:
| Main Authors: | Dou, Yao, Galley, Michel, Peng, Baolin, Kedzie, Chris, Cai, Weixin, Ritter, Alan, Quirk, Chris, Xu, Wei, Gao, Jianfeng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dyna-Think: Synergizing Reasoning, Acting, and World Model Simulation in AI Agents
by: Yu, Xiao, et al.
Published: (2025)
by: Yu, Xiao, et al.
Published: (2025)
Dyna-Mind: Learning to Simulate from Experience for Better AI Agents
by: Yu, Xiao, et al.
Published: (2025)
by: Yu, Xiao, et al.
Published: (2025)
CollabLLM: From Passive Responders to Active Collaborators
by: Wu, Shirley, et al.
Published: (2025)
by: Wu, Shirley, et al.
Published: (2025)
Teaching Language Models to Self-Improve through Interactive Demonstrations
by: Yu, Xiao, et al.
Published: (2023)
by: Yu, Xiao, et al.
Published: (2023)
Self-Checker: Plug-and-Play Modules for Fact-Checking with Large Language Models
by: Li, Miaoran, et al.
Published: (2023)
by: Li, Miaoran, et al.
Published: (2023)
ExACT: Teaching AI Agents to Explore with Reflective-MCTS and Exploratory Learning
by: Yu, Xiao, et al.
Published: (2024)
by: Yu, Xiao, et al.
Published: (2024)
Synthetic Computers at Scale for Long-Horizon Productivity Simulation
by: Ge, Tao, et al.
Published: (2026)
by: Ge, Tao, et al.
Published: (2026)
Lost in Simulation: LLM-Simulated Users are Unreliable Proxies for Human Users in Agentic Evaluations
by: Seshadri, Preethi, et al.
Published: (2026)
by: Seshadri, Preethi, et al.
Published: (2026)
LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language Texts
by: Hashemi, Helia, et al.
Published: (2024)
by: Hashemi, Helia, et al.
Published: (2024)
Measuring and Mitigating the Distributional Gap Between Real and Simulated User Behaviors
by: Mehri, Shuhaib, et al.
Published: (2026)
by: Mehri, Shuhaib, et al.
Published: (2026)
ProPerSim: Developing Proactive and Personalized AI Assistants through User-Assistant Simulation
by: Kim, Jiho, et al.
Published: (2025)
by: Kim, Jiho, et al.
Published: (2025)
Do Androids Know They're Only Dreaming of Electric Sheep?
by: CH-Wang, Sky, et al.
Published: (2023)
by: CH-Wang, Sky, et al.
Published: (2023)
VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation
by: Luo, Ziyang, et al.
Published: (2024)
by: Luo, Ziyang, et al.
Published: (2024)
Synthetic Users, Real Differences: an Evaluation Framework for User Simulation in Multi-Turn Conversations
by: Liu, Yu Lu, et al.
Published: (2026)
by: Liu, Yu Lu, et al.
Published: (2026)
LifeSim: Long-Horizon User Life Simulator for Personalized Assistant Evaluation
by: Duan, Feiyu, et al.
Published: (2026)
by: Duan, Feiyu, et al.
Published: (2026)
Proactive Agent Research Environment: Simulating Active Users to Evaluate Proactive Assistants
by: Nathani, Deepak, et al.
Published: (2026)
by: Nathani, Deepak, et al.
Published: (2026)
Strategic Candidacy in Generative AI Arenas
by: Hays, Chris, et al.
Published: (2026)
by: Hays, Chris, et al.
Published: (2026)
Rethinking Interpretability in the Era of Large Language Models
by: Singh, Chandan, et al.
Published: (2024)
by: Singh, Chandan, et al.
Published: (2024)
Quantifying the Utility of User Simulators for Building Collaborative LLM Assistants
by: Suh, Joseph, et al.
Published: (2026)
by: Suh, Joseph, et al.
Published: (2026)
Separable Expert Architecture: Toward Privacy-Preserving LLM Personalization via Composable Adapters and Deletable User Proxies
by: Schneider, Chris, et al.
Published: (2026)
by: Schneider, Chris, et al.
Published: (2026)
AI User Assistant
by: Sen, Subhayu, et al.
Published: (2025)
by: Sen, Subhayu, et al.
Published: (2025)
Didactic to Constructive: Turning Expert Solutions into Learnable Reasoning
by: Mendes, Ethan, et al.
Published: (2026)
by: Mendes, Ethan, et al.
Published: (2026)
Human– AI partnerships: Living and working with AI Assistants, AI Agents, and AI Companions
by: Ripinka Koli Patil, et al.
Published: (2026)
by: Ripinka Koli Patil, et al.
Published: (2026)
KAUCUS: Knowledge Augmented User Simulators for Training Language Model Assistants
by: Dhole, Kaustubh D.
Published: (2024)
by: Dhole, Kaustubh D.
Published: (2024)
Multi-trait User Simulation with Adaptive Decoding for Conversational Task Assistants
by: Ferreira, Rafael, et al.
Published: (2024)
by: Ferreira, Rafael, et al.
Published: (2024)
"I Can Read but I Can't Turn the Pages."
by: Page, Chris
Published: (1992)
by: Page, Chris
Published: (1992)
AutoSurfer -- Teaching Web Agents through Comprehensive Surfing, Learning, and Modeling
by: Faisal, Fazle Elahi, et al.
Published: (2026)
by: Faisal, Fazle Elahi, et al.
Published: (2026)
ProxyWar: Dynamic Assessment of LLM Code Generation in Game Arenas
by: Peng, Wenjun, et al.
Published: (2026)
by: Peng, Wenjun, et al.
Published: (2026)
Beyond the Assistant Turn: User Turn Generation as a Probe of Interaction Awareness in Language Models
by: Shekkizhar, Sarath, et al.
Published: (2026)
by: Shekkizhar, Sarath, et al.
Published: (2026)
LLM Agent Meets Agentic AI: Can LLM Agents Simulate Customers to Evaluate Agentic-AI-based Shopping Assistants?
by: Sun, Lu, et al.
Published: (2025)
by: Sun, Lu, et al.
Published: (2025)
DriveArena: A Closed-loop Generative Simulation Platform for Autonomous Driving
by: Yang, Xuemeng, et al.
Published: (2024)
by: Yang, Xuemeng, et al.
Published: (2024)
Flexibility and Feedback: A New Approach to Ongoing Training for Reference Student Assistants.
by: Neuhaus, Chris
Published: (2001)
by: Neuhaus, Chris
Published: (2001)
Beyond Benchmarks: How Users Evaluate AI Chat Assistants
by: Awan, Moiz Sadiq, et al.
Published: (2026)
by: Awan, Moiz Sadiq, et al.
Published: (2026)
Simulation Based Composite Likelihood
by: Rimella, Lorenzo, et al.
Published: (2023)
by: Rimella, Lorenzo, et al.
Published: (2023)
Reliable Simulation of Quantum Channels: the Error Exponent
by: Li, Ke, et al.
Published: (2021)
by: Li, Ke, et al.
Published: (2021)
Towards Decoding Developer Cognition in the Age of AI Assistants
by: Haque, Ebtesam Al, et al.
Published: (2025)
by: Haque, Ebtesam Al, et al.
Published: (2025)
AgentEval: Generative Agents as Reliable Proxies for Human Evaluation of AI-Generated Content
by: Vu, Thanh, et al.
Published: (2025)
by: Vu, Thanh, et al.
Published: (2025)
Test-Time Learning with an Evolving Library
by: Xu, Weijia, et al.
Published: (2026)
by: Xu, Weijia, et al.
Published: (2026)
User Simulation in the Era of Generative AI: User Modeling, Synthetic Data Generation, and System Evaluation
by: Balog, Krisztian, et al.
Published: (2025)
by: Balog, Krisztian, et al.
Published: (2025)
How Reliable is Your Simulator? Analysis on the Limitations of Current LLM-based User Simulators for Conversational Recommendation
by: Zhu, Lixi, et al.
Published: (2024)
by: Zhu, Lixi, et al.
Published: (2024)
Similar Items
-
Dyna-Think: Synergizing Reasoning, Acting, and World Model Simulation in AI Agents
by: Yu, Xiao, et al.
Published: (2025) -
Dyna-Mind: Learning to Simulate from Experience for Better AI Agents
by: Yu, Xiao, et al.
Published: (2025) -
CollabLLM: From Passive Responders to Active Collaborators
by: Wu, Shirley, et al.
Published: (2025) -
Teaching Language Models to Self-Improve through Interactive Demonstrations
by: Yu, Xiao, et al.
Published: (2023) -
Self-Checker: Plug-and-Play Modules for Fact-Checking with Large Language Models
by: Li, Miaoran, et al.
Published: (2023)