Lost in Simulation: LLM-Simulated Users are Unreliable Proxies for Human Users in Agentic Evaluations
Fuente:
arXiv
Saved in:
| Main Authors: | Seshadri, Preethi, Cahyawijaya, Samuel, Odumakinde, Ayomide, Singh, Sameer, Goldfarb-Tarrant, Seraphina |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Imperfectly Cooperative Human-AI Interactions: Comparing the Impacts of Human and AI Attributes in Simulated and User Studies
by: Cohen, Myke C., et al.
Published: (2026)
by: Cohen, Myke C., et al.
Published: (2026)
Who Defines "Best"? Towards Interactive, User-Defined Evaluation of LLM Leaderboards
by: Jung, Minji, et al.
Published: (2026)
by: Jung, Minji, et al.
Published: (2026)
Mirroring Users: Towards Building Preference-aligned User Simulator with User Feedback in Recommendation
by: Wei, Tianjun, et al.
Published: (2025)
by: Wei, Tianjun, et al.
Published: (2025)
Assessing the Human-Likeness of LLM-Driven Digital Twins in Simulating Health Care System Trust
by: Wu, Yuzhou, et al.
Published: (2025)
by: Wu, Yuzhou, et al.
Published: (2025)
Generative AI User Experience: Developing Human--AI Epistemic Partnership
by: Zhai, Xiaoming
Published: (2026)
by: Zhai, Xiaoming
Published: (2026)
Avenir-UX: Automated UX Evaluation via Simulated Human Web Interaction with GUI Grounding
by: Tan, Wee Joe, et al.
Published: (2026)
by: Tan, Wee Joe, et al.
Published: (2026)
Small Changes, Large Consequences: Analyzing the Allocational Fairness of LLMs in Hiring Contexts
by: Seshadri, Preethi, et al.
Published: (2025)
by: Seshadri, Preethi, et al.
Published: (2025)
Exploring User Acceptance and Concerns toward LLM-powered Conversational Agents in Immersive Extended Reality
by: Bozkir, Efe, et al.
Published: (2025)
by: Bozkir, Efe, et al.
Published: (2025)
AgentSUMO: An Agentic Framework for Interactive Simulation Scenario Generation in SUMO via Large Language Models
by: Jeong, Minwoo, et al.
Published: (2025)
by: Jeong, Minwoo, et al.
Published: (2025)
Agentic Persona Control and Task State Tracking for Realistic User Simulation in Interactive Scenarios
by: Karthikeyan, Hareeshwar
Published: (2025)
by: Karthikeyan, Hareeshwar
Published: (2025)
RecUserSim: A Realistic and Diverse User Simulator for Evaluating Conversational Recommender Systems
by: Chen, Luyu, et al.
Published: (2025)
by: Chen, Luyu, et al.
Published: (2025)
A LLM-based Controllable, Scalable, Human-Involved User Simulator Framework for Conversational Recommender Systems
by: Zhu, Lixi, et al.
Published: (2024)
by: Zhu, Lixi, et al.
Published: (2024)
The Impacts of AI Avatar Appearance and Disclosure on User Motivation
by: Visser, Boele, et al.
Published: (2024)
by: Visser, Boele, et al.
Published: (2024)
LLM Social Simulations Are a Promising Research Method
by: Anthis, Jacy Reese, et al.
Published: (2025)
by: Anthis, Jacy Reese, et al.
Published: (2025)
Investigating Thematic Patterns and User Preferences in LLM Interactions using BERTopic
by: Bhandarkar, Abhay, et al.
Published: (2025)
by: Bhandarkar, Abhay, et al.
Published: (2025)
A Checklist for Trustworthy, Safe, and User-Friendly Mental Health Chatbots
by: Haran, Shreya, et al.
Published: (2026)
by: Haran, Shreya, et al.
Published: (2026)
User Simulation for Evaluating Information Access Systems
by: Balog, Krisztian, et al.
Published: (2023)
by: Balog, Krisztian, et al.
Published: (2023)
Scalable LLM-based Coding of Dialogue in Healthcare Simulation: Balancing Coding Performance, Processing Time, and Environmental Impact
by: Garces, Kiyoshige, et al.
Published: (2026)
by: Garces, Kiyoshige, et al.
Published: (2026)
Malicious LLM-Based Conversational AI Makes Users Reveal Personal Information
by: Zhan, Xiao, et al.
Published: (2025)
by: Zhan, Xiao, et al.
Published: (2025)
Towards User-Centred Design of AI-Assisted Decision-Making in Law Enforcement
by: Nowack, Vesna, et al.
Published: (2025)
by: Nowack, Vesna, et al.
Published: (2025)
What Is Required for Empathic AI? It Depends, and Why That Matters for AI Developers and Users
by: Borg, Jana Schaich, et al.
Published: (2024)
by: Borg, Jana Schaich, et al.
Published: (2024)
Agentic Enterprise: AI-Centric User to User-Centric AI
by: Narechania, Arpit, et al.
Published: (2025)
by: Narechania, Arpit, et al.
Published: (2025)
Human Control Is the Anchor, Not the Answer: Early Divergence of Oversight in Agentic AI Communities
by: Shi, Hanjing, et al.
Published: (2026)
by: Shi, Hanjing, et al.
Published: (2026)
The Imbalanced User-AI Relationships as an Ethical Failure of Front-End Design in Healthcare AI
by: Mwadime, Maureen Mghambi
Published: (2026)
by: Mwadime, Maureen Mghambi
Published: (2026)
From Sea to System: Exploring User-Centered Explainable AI for Maritime Decision Support
by: Jirak, Doreen, et al.
Published: (2025)
by: Jirak, Doreen, et al.
Published: (2025)
User Perceptions vs. Proxy LLM Judges: Privacy and Helpfulness in LLM Responses to Privacy-Sensitive Scenarios
by: Wu, Xiaoyuan, et al.
Published: (2025)
by: Wu, Xiaoyuan, et al.
Published: (2025)
Emoji Reactions on Telegram: Unreliable Indicators of Emotional Resonance
by: Tardelli, Serena, et al.
Published: (2025)
by: Tardelli, Serena, et al.
Published: (2025)
AI-Driven Feedback Loops in Digital Technologies: Psychological Impacts on User Behaviour and Well-Being
by: Adanyin, Anthonette
Published: (2024)
by: Adanyin, Anthonette
Published: (2024)
User Negotiations of Authenticity, Ownership, and Governance on AI-Generated Video Platforms: Evidence from Sora
by: Shen, Bohui, et al.
Published: (2025)
by: Shen, Bohui, et al.
Published: (2025)
PREFINE: Personalized Story Generation via Simulated User Critics and User-Specific Rubric Generation
by: Ueda, Kentaro, et al.
Published: (2025)
by: Ueda, Kentaro, et al.
Published: (2025)
RealUserSim: Bridging the Reality Gap in Agent Benchmarking via Grounded User Simulation
by: Zhu, Ming, et al.
Published: (2026)
by: Zhu, Ming, et al.
Published: (2026)
Who's Asking? Simulating Role-Based Questions for Conversational AI Evaluation
by: Kaur, Navreet, et al.
Published: (2025)
by: Kaur, Navreet, et al.
Published: (2025)
Agentic AI as Undercover Teammates: Argumentative Knowledge Construction in Hybrid Human-AI Collaborative Learning
by: Yan, Lixiang, et al.
Published: (2025)
by: Yan, Lixiang, et al.
Published: (2025)
MetaScoreLens: Evaluating User Feedback Across Digital Entertainment Systems
by: Ellington, Christian, et al.
Published: (2025)
by: Ellington, Christian, et al.
Published: (2025)
User Intent to Use DeepSeek for Healthcare Purposes and their Trust in the Large Language Model: Multinational Survey Study
by: Choudhury, Avishek, et al.
Published: (2025)
by: Choudhury, Avishek, et al.
Published: (2025)
Confident Teacher, Confident Student? A Novel User Study Design for Investigating the Didactic Potential of Explanations and their Impact on Uncertainty
by: Chiaburu, Teodor, et al.
Published: (2024)
by: Chiaburu, Teodor, et al.
Published: (2024)
Evaluating Multimodal Language Models as Visual Assistants for Visually Impaired Users
by: Karamolegkou, Antonia, et al.
Published: (2025)
by: Karamolegkou, Antonia, et al.
Published: (2025)
Measuring User Perceived Security of Mobile Banking Applications
by: Apaua, Richard, et al.
Published: (2022)
by: Apaua, Richard, et al.
Published: (2022)
ChatGPT and U(X): A Rapid Review on Measuring the User Experience
by: Seaborn, Katie
Published: (2025)
by: Seaborn, Katie
Published: (2025)
Llms, Virtual Users, and Bias: Predicting Any Survey Question Without Human Data
by: Sinacola, Enzo, et al.
Published: (2025)
by: Sinacola, Enzo, et al.
Published: (2025)
Similar Items
-
Imperfectly Cooperative Human-AI Interactions: Comparing the Impacts of Human and AI Attributes in Simulated and User Studies
by: Cohen, Myke C., et al.
Published: (2026) -
Who Defines "Best"? Towards Interactive, User-Defined Evaluation of LLM Leaderboards
by: Jung, Minji, et al.
Published: (2026) -
Mirroring Users: Towards Building Preference-aligned User Simulator with User Feedback in Recommendation
by: Wei, Tianjun, et al.
Published: (2025) -
Assessing the Human-Likeness of LLM-Driven Digital Twins in Simulating Health Care System Trust
by: Wu, Yuzhou, et al.
Published: (2025) -
Generative AI User Experience: Developing Human--AI Epistemic Partnership
by: Zhai, Xiaoming
Published: (2026)