Quantifying the Utility of User Simulators for Building Collaborative LLM Assistants
Fuente:
arXiv
Guardado en:
| Autores principales: | Suh, Joseph, Raj, Ayush, Kang, Minwoo, Chang, Serina |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Language Model Fine-Tuning on Scaled Survey Data for Predicting Distributions of Public Opinions
por: Suh, Joseph, et al.
Publicado: (2025)
por: Suh, Joseph, et al.
Publicado: (2025)
Graph-Based Alternatives to LLMs for Human Simulation
por: Suh, Joseph, et al.
Publicado: (2025)
por: Suh, Joseph, et al.
Publicado: (2025)
Deep Binding of Language Model Virtual Personas: a Study on Approximating Political Partisan Misperceptions
por: Kang, Minwoo, et al.
Publicado: (2025)
por: Kang, Minwoo, et al.
Publicado: (2025)
Identity, Cooperation and Framing Effects within Groups of Real and Simulated Humans
por: Moon, Suhong, et al.
Publicado: (2026)
por: Moon, Suhong, et al.
Publicado: (2026)
Rediscovering the Latent Dimensions of Personality with Large Language Models as Trait Descriptors
por: Suh, Joseph, et al.
Publicado: (2024)
por: Suh, Joseph, et al.
Publicado: (2024)
Valid Survey Simulations with Limited Human Data: The Roles of Prompting, Fine-Tuning, and Rectification
por: Krsteski, Stefan, et al.
Publicado: (2025)
por: Krsteski, Stefan, et al.
Publicado: (2025)
ProPerSim: Developing Proactive and Personalized AI Assistants through User-Assistant Simulation
por: Kim, Jiho, et al.
Publicado: (2025)
por: Kim, Jiho, et al.
Publicado: (2025)
SimulatorArena: Are User Simulators Reliable Proxies for Multi-Turn Evaluation of AI Assistants?
por: Dou, Yao, et al.
Publicado: (2025)
por: Dou, Yao, et al.
Publicado: (2025)
Quantifying the Persona Effect in LLM Simulations
por: Hu, Tiancheng, et al.
Publicado: (2024)
por: Hu, Tiancheng, et al.
Publicado: (2024)
LifeSim: Long-Horizon User Life Simulator for Personalized Assistant Evaluation
por: Duan, Feiyu, et al.
Publicado: (2026)
por: Duan, Feiyu, et al.
Publicado: (2026)
Non-Collaborative User Simulators for Tool Agents
por: Shim, Jeonghoon, et al.
Publicado: (2025)
por: Shim, Jeonghoon, et al.
Publicado: (2025)
SafeChat: A Framework for Building Trustworthy Collaborative Assistants and a Case Study of its Usefulness
por: Srivastava, Biplav, et al.
Publicado: (2025)
por: Srivastava, Biplav, et al.
Publicado: (2025)
ChatBench: From Static Benchmarks to Human-AI Evaluation
por: Chang, Serina, et al.
Publicado: (2025)
por: Chang, Serina, et al.
Publicado: (2025)
LLM-based Medical Assistant Personalization with Short- and Long-Term Memory Coordination
por: Zhang, Kai, et al.
Publicado: (2023)
por: Zhang, Kai, et al.
Publicado: (2023)
User-Assistant Bias in LLMs
por: Pan, Xu, et al.
Publicado: (2025)
por: Pan, Xu, et al.
Publicado: (2025)
Virtual Personas for Language Models via an Anthology of Backstories
por: Moon, Suhong, et al.
Publicado: (2024)
por: Moon, Suhong, et al.
Publicado: (2024)
When Wording Steers the Evaluation: Framing Bias in LLM judges
por: Hwang, Yerin, et al.
Publicado: (2026)
por: Hwang, Yerin, et al.
Publicado: (2026)
Multi-trait User Simulation with Adaptive Decoding for Conversational Task Assistants
por: Ferreira, Rafael, et al.
Publicado: (2024)
por: Ferreira, Rafael, et al.
Publicado: (2024)
Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation
por: Lee, Dongryeol, et al.
Publicado: (2026)
por: Lee, Dongryeol, et al.
Publicado: (2026)
User Modeling Challenges in Interactive AI Assistant Systems
por: Su, Megan, et al.
Publicado: (2024)
por: Su, Megan, et al.
Publicado: (2024)
MathVC: An LLM-Simulated Multi-Character Virtual Classroom for Mathematics Education
por: Yue, Murong, et al.
Publicado: (2024)
por: Yue, Murong, et al.
Publicado: (2024)
An EcoSage Assistant: Towards Building A Multimodal Plant Care Dialogue Assistant
por: Tomar, Mohit, et al.
Publicado: (2024)
por: Tomar, Mohit, et al.
Publicado: (2024)
SafeSearch: Do Not Trade Safety for Utility in LLM Search Agents
por: Zhan, Qiusi, et al.
Publicado: (2025)
por: Zhan, Qiusi, et al.
Publicado: (2025)
A Survey on LLM-based Conversational User Simulation
por: Ni, Bo, et al.
Publicado: (2026)
por: Ni, Bo, et al.
Publicado: (2026)
Human-AI Collaborative Taxonomy Construction: A Case Study in Profession-Specific Writing Assistants
por: Lee, Minhwa, et al.
Publicado: (2024)
por: Lee, Minhwa, et al.
Publicado: (2024)
Building A Coding Assistant via the Retrieval-Augmented Language Model
por: Li, Xinze, et al.
Publicado: (2024)
por: Li, Xinze, et al.
Publicado: (2024)
MemSim: A Bayesian Simulator for Evaluating Memory of LLM-based Personal Assistants
por: Zhang, Zeyu, et al.
Publicado: (2024)
por: Zhang, Zeyu, et al.
Publicado: (2024)
Improving LLM-Powered EDA Assistants with RAFT
por: Shi, Luyao, et al.
Publicado: (2025)
por: Shi, Luyao, et al.
Publicado: (2025)
CARE: A Clue-guided Assistant for CSRs to Read User Manuals
por: Du, Weihong, et al.
Publicado: (2024)
por: Du, Weihong, et al.
Publicado: (2024)
LLM Agent Meets Agentic AI: Can LLM Agents Simulate Customers to Evaluate Agentic-AI-based Shopping Assistants?
por: Sun, Lu, et al.
Publicado: (2025)
por: Sun, Lu, et al.
Publicado: (2025)
Plan-Then-Execute: An Empirical Study of User Trust and Team Performance When Using LLM Agents As A Daily Assistant
por: He, Gaole, et al.
Publicado: (2025)
por: He, Gaole, et al.
Publicado: (2025)
Reliable LLM-based User Simulator for Task-Oriented Dialogue Systems
por: Sekulić, Ivan, et al.
Publicado: (2024)
por: Sekulić, Ivan, et al.
Publicado: (2024)
KAUCUS: Knowledge Augmented User Simulators for Training Language Model Assistants
por: Dhole, Kaustubh D.
Publicado: (2024)
por: Dhole, Kaustubh D.
Publicado: (2024)
Collaborating Action by Action: A Multi-agent LLM Framework for Embodied Reasoning
por: White, Isadora, et al.
Publicado: (2025)
por: White, Isadora, et al.
Publicado: (2025)
Personality Matters: User Traits Predict LLM Preferences in Multi-Turn Collaborative Tasks
por: Yunusov, Sarfaroz, et al.
Publicado: (2025)
por: Yunusov, Sarfaroz, et al.
Publicado: (2025)
Goal Alignment in LLM-Based User Simulators for Conversational AI
por: Mehri, Shuhaib, et al.
Publicado: (2025)
por: Mehri, Shuhaib, et al.
Publicado: (2025)
Users Mispredict Their Own Preferences for AI Writing Assistance
por: Lai, Vivian, et al.
Publicado: (2026)
por: Lai, Vivian, et al.
Publicado: (2026)
Counterfactual Fairness Evaluation of LLM-Based Contact Center Agent Quality Assurance System
por: Mayilvaghanan, Kawin, et al.
Publicado: (2026)
por: Mayilvaghanan, Kawin, et al.
Publicado: (2026)
DuetSim: Building User Simulator with Dual Large Language Models for Task-Oriented Dialogues
por: Luo, Xiang, et al.
Publicado: (2024)
por: Luo, Xiang, et al.
Publicado: (2024)
From AI Assistant to AI Scientist: Autonomous Discovery of LLM-RL Algorithms with LLM Agents
por: Xia, Sirui, et al.
Publicado: (2026)
por: Xia, Sirui, et al.
Publicado: (2026)
Ejemplares similares
-
Language Model Fine-Tuning on Scaled Survey Data for Predicting Distributions of Public Opinions
por: Suh, Joseph, et al.
Publicado: (2025) -
Graph-Based Alternatives to LLMs for Human Simulation
por: Suh, Joseph, et al.
Publicado: (2025) -
Deep Binding of Language Model Virtual Personas: a Study on Approximating Political Partisan Misperceptions
por: Kang, Minwoo, et al.
Publicado: (2025) -
Identity, Cooperation and Framing Effects within Groups of Real and Simulated Humans
por: Moon, Suhong, et al.
Publicado: (2026) -
Rediscovering the Latent Dimensions of Personality with Large Language Models as Trait Descriptors
por: Suh, Joseph, et al.
Publicado: (2024)