Mind the Sim2Real Gap in User Simulation for Agentic Tasks
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhou, Xuhui, Sun, Weiwei, Ma, Qianou, Xie, Yiqing, Liu, Jiarui, Du, Weihua, Welleck, Sean, Yang, Yiming, Neubig, Graham, Wu, Sherry Tongshuang, Sap, Maarten |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Reinforcing Human Behavior Simulation via Verbal Feedback
di: Sun, Weiwei, et al.
Pubblicazione: (2026)
di: Sun, Weiwei, et al.
Pubblicazione: (2026)
Training Proactive and Personalized LLM Agents
di: Sun, Weiwei, et al.
Pubblicazione: (2025)
di: Sun, Weiwei, et al.
Pubblicazione: (2025)
Agentic-R1: Distilled Dual-Strategy Reasoning
di: Du, Weihua, et al.
Pubblicazione: (2025)
di: Du, Weihua, et al.
Pubblicazione: (2025)
Optimizing Temperature for Language Models with Multi-Sample Inference
di: Du, Weihua, et al.
Pubblicazione: (2025)
di: Du, Weihua, et al.
Pubblicazione: (2025)
TOM-SWE: User Mental Modeling For Software Engineering Agents
di: Zhou, Xuhui, et al.
Pubblicazione: (2025)
di: Zhou, Xuhui, et al.
Pubblicazione: (2025)
Ambig-SWE: Interactive Agents to Overcome Underspecificity in Software Engineering
di: Vijayvargiya, Sanidhya, et al.
Pubblicazione: (2025)
di: Vijayvargiya, Sanidhya, et al.
Pubblicazione: (2025)
Gym-Anything: Turn any Software into an Agent Environment
di: Aggarwal, Pranjal, et al.
Pubblicazione: (2026)
di: Aggarwal, Pranjal, et al.
Pubblicazione: (2026)
GradAlign: Gradient-Aligned Data Selection for LLM Reinforcement Learning
di: Yang, Ningyuan, et al.
Pubblicazione: (2026)
di: Yang, Ningyuan, et al.
Pubblicazione: (2026)
Stereotype or Personalization? User Identity Biases Chatbot Recommendations
di: Kantharuban, Anjali, et al.
Pubblicazione: (2024)
di: Kantharuban, Anjali, et al.
Pubblicazione: (2024)
Rethinking Theory of Mind Benchmarks for LLMs: Towards A User-Centered Perspective
di: Wang, Qiaosi, et al.
Pubblicazione: (2025)
di: Wang, Qiaosi, et al.
Pubblicazione: (2025)
OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety
di: Vijayvargiya, Sanidhya, et al.
Pubblicazione: (2025)
di: Vijayvargiya, Sanidhya, et al.
Pubblicazione: (2025)
Not Everyone Wins with LLMs: Behavioral Patterns and Pedagogical Implications for AI Literacy in Programmatic Data Science
di: Ma, Qianou, et al.
Pubblicazione: (2025)
di: Ma, Qianou, et al.
Pubblicazione: (2025)
Do LLMs exhibit human-like response biases? A case study in survey design
di: Tjuatja, Lindia, et al.
Pubblicazione: (2023)
di: Tjuatja, Lindia, et al.
Pubblicazione: (2023)
Social World Models
di: Zhou, Xuhui, et al.
Pubblicazione: (2025)
di: Zhou, Xuhui, et al.
Pubblicazione: (2025)
SELF-GUIDE: Better Task-Specific Instruction Following via Self-Synthetic Finetuning
di: Zhao, Chenyang, et al.
Pubblicazione: (2024)
di: Zhao, Chenyang, et al.
Pubblicazione: (2024)
How to Teach Programming in the AI Era? Using LLMs as a Teachable Agent for Debugging
di: Ma, Qianou, et al.
Pubblicazione: (2023)
di: Ma, Qianou, et al.
Pubblicazione: (2023)
SoMi-ToM: Evaluating Multi-Perspective Theory of Mind in Embodied Social Interactions
di: Fan, Xianzhe, et al.
Pubblicazione: (2025)
di: Fan, Xianzhe, et al.
Pubblicazione: (2025)
Is this the real life? Is this just fantasy? The Misleading Success of Simulating Social Interactions With LLMs
di: Zhou, Xuhui, et al.
Pubblicazione: (2024)
di: Zhou, Xuhui, et al.
Pubblicazione: (2024)
BIG5-CHAT: Shaping LLM Personalities Through Training on Human-Grounded Data
di: Li, Wenkai, et al.
Pubblicazione: (2024)
di: Li, Wenkai, et al.
Pubblicazione: (2024)
Inducing Programmatic Skills for Agentic Tasks
di: Wang, Zora Zhiruo, et al.
Pubblicazione: (2025)
di: Wang, Zora Zhiruo, et al.
Pubblicazione: (2025)
SOTOPIA-TOM: Evaluating Information Management in Multi-Agent Interaction with Theory of Mind
di: YS, Yashwanth, et al.
Pubblicazione: (2026)
di: YS, Yashwanth, et al.
Pubblicazione: (2026)
RealUserSim: Bridging the Reality Gap in Agent Benchmarking via Grounded User Simulation
di: Zhu, Ming, et al.
Pubblicazione: (2026)
di: Zhu, Ming, et al.
Pubblicazione: (2026)
User-Driven Value Alignment: Understanding Users' Perceptions and Strategies for Addressing Biased and Discriminatory Statements in AI Companions
di: Fan, Xianzhe, et al.
Pubblicazione: (2024)
di: Fan, Xianzhe, et al.
Pubblicazione: (2024)
Mind the Sim-to-Real Gap & Think Like a Scientist
di: Parikh, Harsh, et al.
Pubblicazione: (2026)
di: Parikh, Harsh, et al.
Pubblicazione: (2026)
RECAP: An End-to-End Platform for Capturing, Replaying, and Analyzing AI-Assisted Programming Interactions
di: He, Keyu, et al.
Pubblicazione: (2026)
di: He, Keyu, et al.
Pubblicazione: (2026)
From Decoding to Meta-Generation: Inference-time Algorithms for Large Language Models
di: Welleck, Sean, et al.
Pubblicazione: (2024)
di: Welleck, Sean, et al.
Pubblicazione: (2024)
Better Synthetic Data by Retrieving and Transforming Existing Datasets
di: Gandhi, Saumya, et al.
Pubblicazione: (2024)
di: Gandhi, Saumya, et al.
Pubblicazione: (2024)
Mind the Gesture: Evaluating AI Sensitivity to Culturally Offensive Non-Verbal Gestures
di: Yerukola, Akhila, et al.
Pubblicazione: (2025)
di: Yerukola, Akhila, et al.
Pubblicazione: (2025)
Minion: A Technology Probe to Explore How Users Negotiate Harmful Value Conflicts with AI Companions
di: Fan, Xianzhe, et al.
Pubblicazione: (2024)
di: Fan, Xianzhe, et al.
Pubblicazione: (2024)
Imperfectly Cooperative Human-AI Interactions: Comparing the Impacts of Human and AI Attributes in Simulated and User Studies
di: Cohen, Myke C., et al.
Pubblicazione: (2026)
di: Cohen, Myke C., et al.
Pubblicazione: (2026)
1-2-3 Check: Enhancing Contextual Privacy in LLM via Multi-Agent Reasoning
di: Li, Wenkai, et al.
Pubblicazione: (2025)
di: Li, Wenkai, et al.
Pubblicazione: (2025)
GoodPoint: Learning Constructive Scientific Paper Feedback from Author Responses
di: Mun, Jimin, et al.
Pubblicazione: (2026)
di: Mun, Jimin, et al.
Pubblicazione: (2026)
What Should We Engineer in Prompts? Training Humans in Requirement-Driven LLM Use
di: Ma, Qianou, et al.
Pubblicazione: (2024)
di: Ma, Qianou, et al.
Pubblicazione: (2024)
SOTOPIA-S4: a user-friendly system for flexible, customizable, and large-scale social simulation
di: Zhou, Xuhui, et al.
Pubblicazione: (2025)
di: Zhou, Xuhui, et al.
Pubblicazione: (2025)
Lean-STaR: Learning to Interleave Thinking and Proving
di: Lin, Haohan, et al.
Pubblicazione: (2024)
di: Lin, Haohan, et al.
Pubblicazione: (2024)
Quantification of Sim2Real Gap via Neural Simulation Gap Function
di: Sangeerth, P, et al.
Pubblicazione: (2025)
di: Sangeerth, P, et al.
Pubblicazione: (2025)
SOTOPIA-$π$: Interactive Learning of Socially Intelligent Language Agents
di: Wang, Ruiyi, et al.
Pubblicazione: (2024)
di: Wang, Ruiyi, et al.
Pubblicazione: (2024)
SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents
di: Zhou, Xuhui, et al.
Pubblicazione: (2023)
di: Zhou, Xuhui, et al.
Pubblicazione: (2023)
Learning to Dock: A Simulation-based Study on Closing the Sim2Real Gap in Autonomous Underwater Docking
di: Chang, Kevin, et al.
Pubblicazione: (2025)
di: Chang, Kevin, et al.
Pubblicazione: (2025)
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
di: Aggarwal, Pranjal, et al.
Pubblicazione: (2025)
di: Aggarwal, Pranjal, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Reinforcing Human Behavior Simulation via Verbal Feedback
di: Sun, Weiwei, et al.
Pubblicazione: (2026) -
Training Proactive and Personalized LLM Agents
di: Sun, Weiwei, et al.
Pubblicazione: (2025) -
Agentic-R1: Distilled Dual-Strategy Reasoning
di: Du, Weihua, et al.
Pubblicazione: (2025) -
Optimizing Temperature for Language Models with Multi-Sample Inference
di: Du, Weihua, et al.
Pubblicazione: (2025) -
TOM-SWE: User Mental Modeling For Software Engineering Agents
di: Zhou, Xuhui, et al.
Pubblicazione: (2025)