Reinforcing Human Behavior Simulation via Verbal Feedback
Fuente:
arXiv
Saved in:
| Main Authors: | Sun, Weiwei, Zhou, Xuhui, Liu, Jiarui, Du, Weihua, Sun, Haojia, Xie, Yiqing, Ma, Qianou, Chen, Sihao, Wan, Mengting, Yang, Longqi, Zhou, Pei, Wu, Sherry, Welleck, Sean, Neubig, Graham, Yang, Yiming, Sap, Maarten |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mind the Sim2Real Gap in User Simulation for Agentic Tasks
by: Zhou, Xuhui, et al.
Published: (2026)
by: Zhou, Xuhui, et al.
Published: (2026)
Training Proactive and Personalized LLM Agents
by: Sun, Weiwei, et al.
Published: (2025)
by: Sun, Weiwei, et al.
Published: (2025)
GradAlign: Gradient-Aligned Data Selection for LLM Reinforcement Learning
by: Yang, Ningyuan, et al.
Published: (2026)
by: Yang, Ningyuan, et al.
Published: (2026)
Ambig-SWE: Interactive Agents to Overcome Underspecificity in Software Engineering
by: Vijayvargiya, Sanidhya, et al.
Published: (2025)
by: Vijayvargiya, Sanidhya, et al.
Published: (2025)
Optimizing Temperature for Language Models with Multi-Sample Inference
by: Du, Weihua, et al.
Published: (2025)
by: Du, Weihua, et al.
Published: (2025)
Gym-Anything: Turn any Software into an Agent Environment
by: Aggarwal, Pranjal, et al.
Published: (2026)
by: Aggarwal, Pranjal, et al.
Published: (2026)
TOM-SWE: User Mental Modeling For Software Engineering Agents
by: Zhou, Xuhui, et al.
Published: (2025)
by: Zhou, Xuhui, et al.
Published: (2025)
Teaching Language Models To Gather Information Proactively
by: Huang, Tenghao, et al.
Published: (2025)
by: Huang, Tenghao, et al.
Published: (2025)
Agentic-R1: Distilled Dual-Strategy Reasoning
by: Du, Weihua, et al.
Published: (2025)
by: Du, Weihua, et al.
Published: (2025)
Social World Models
by: Zhou, Xuhui, et al.
Published: (2025)
by: Zhou, Xuhui, et al.
Published: (2025)
GoodPoint: Learning Constructive Scientific Paper Feedback from Author Responses
by: Mun, Jimin, et al.
Published: (2026)
by: Mun, Jimin, et al.
Published: (2026)
OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety
by: Vijayvargiya, Sanidhya, et al.
Published: (2025)
by: Vijayvargiya, Sanidhya, et al.
Published: (2025)
1-2-3 Check: Enhancing Contextual Privacy in LLM via Multi-Agent Reasoning
by: Li, Wenkai, et al.
Published: (2025)
by: Li, Wenkai, et al.
Published: (2025)
Lean-STaR: Learning to Interleave Thinking and Proving
by: Lin, Haohan, et al.
Published: (2024)
by: Lin, Haohan, et al.
Published: (2024)
BIG5-CHAT: Shaping LLM Personalities Through Training on Human-Grounded Data
by: Li, Wenkai, et al.
Published: (2024)
by: Li, Wenkai, et al.
Published: (2024)
Evaluating LLM-Simulated Conversations in Modeling Inconsistent and Uncollaborative Behaviors in Human Social Interaction
by: Kamoi, Ryo, et al.
Published: (2026)
by: Kamoi, Ryo, et al.
Published: (2026)
Stereotype or Personalization? User Identity Biases Chatbot Recommendations
by: Kantharuban, Anjali, et al.
Published: (2024)
by: Kantharuban, Anjali, et al.
Published: (2024)
WildFeedback: Aligning LLMs With In-situ User Interactions And Feedback
by: Shi, Taiwei, et al.
Published: (2024)
by: Shi, Taiwei, et al.
Published: (2024)
Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models
by: Wu, Yangzhen, et al.
Published: (2024)
by: Wu, Yangzhen, et al.
Published: (2024)
Is this the real life? Is this just fantasy? The Misleading Success of Simulating Social Interactions With LLMs
by: Zhou, Xuhui, et al.
Published: (2024)
by: Zhou, Xuhui, et al.
Published: (2024)
Rethinking Theory of Mind Benchmarks for LLMs: Towards A User-Centered Perspective
by: Wang, Qiaosi, et al.
Published: (2025)
by: Wang, Qiaosi, et al.
Published: (2025)
Mind the Gesture: Evaluating AI Sensitivity to Culturally Offensive Non-Verbal Gestures
by: Yerukola, Akhila, et al.
Published: (2025)
by: Yerukola, Akhila, et al.
Published: (2025)
User-Driven Value Alignment: Understanding Users' Perceptions and Strategies for Addressing Biased and Discriminatory Statements in AI Companions
by: Fan, Xianzhe, et al.
Published: (2024)
by: Fan, Xianzhe, et al.
Published: (2024)
Beyond Output Critique: Self-Correction via Task Distillation
by: Rahmani, Hossein A., et al.
Published: (2026)
by: Rahmani, Hossein A., et al.
Published: (2026)
SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents
by: Zhou, Xuhui, et al.
Published: (2023)
by: Zhou, Xuhui, et al.
Published: (2023)
SoMi-ToM: Evaluating Multi-Perspective Theory of Mind in Embodied Social Interactions
by: Fan, Xianzhe, et al.
Published: (2025)
by: Fan, Xianzhe, et al.
Published: (2025)
PolygloToxicityPrompts: Multilingual Evaluation of Neural Toxic Degeneration in Large Language Models
by: Jain, Devansh, et al.
Published: (2024)
by: Jain, Devansh, et al.
Published: (2024)
AutoPresent: Designing Structured Visuals from Scratch
by: Ge, Jiaxin, et al.
Published: (2025)
by: Ge, Jiaxin, et al.
Published: (2025)
Words Like Knives: Backstory-Personalized Modeling and Detection of Violent Communication
by: Shen, Jocelyn, et al.
Published: (2025)
by: Shen, Jocelyn, et al.
Published: (2025)
Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision
by: Sun, Zhiqing, et al.
Published: (2024)
by: Sun, Zhiqing, et al.
Published: (2024)
Corporate Communication Companion (CCC): An LLM-empowered Writing Assistant for Workplace Social Media
by: Lu, Zhuoran, et al.
Published: (2024)
by: Lu, Zhuoran, et al.
Published: (2024)
GenTool: Enhancing Tool Generalization in Language Models through Zero-to-One and Weak-to-Strong Simulation
by: He, Jie, et al.
Published: (2025)
by: He, Jie, et al.
Published: (2025)
AdaExplore: Failure-Driven Adaptation and Diversity-Preserving Search for Efficient Kernel Generation
by: Du, Weihua, et al.
Published: (2026)
by: Du, Weihua, et al.
Published: (2026)
One Model, All Roles: Multi-Turn, Multi-Agent Self-Play Reinforcement Learning for Conversational Social Intelligence
by: Jiang, Bowen, et al.
Published: (2026)
by: Jiang, Bowen, et al.
Published: (2026)
From Decoding to Meta-Generation: Inference-time Algorithms for Large Language Models
by: Welleck, Sean, et al.
Published: (2024)
by: Welleck, Sean, et al.
Published: (2024)
SOTOPIA-TOM: Evaluating Information Management in Multi-Agent Interaction with Theory of Mind
by: YS, Yashwanth, et al.
Published: (2026)
by: YS, Yashwanth, et al.
Published: (2026)
AI-LieDar: Examine the Trade-off Between Utility and Truthfulness in LLM Agents
by: Su, Zhe, et al.
Published: (2024)
by: Su, Zhe, et al.
Published: (2024)
Minion: A Technology Probe to Explore How Users Negotiate Harmful Value Conflicts with AI Companions
by: Fan, Xianzhe, et al.
Published: (2024)
by: Fan, Xianzhe, et al.
Published: (2024)
Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory
by: Mireshghallah, Niloofar, et al.
Published: (2023)
by: Mireshghallah, Niloofar, et al.
Published: (2023)
SOTOPIA-$π$: Interactive Learning of Socially Intelligent Language Agents
by: Wang, Ruiyi, et al.
Published: (2024)
by: Wang, Ruiyi, et al.
Published: (2024)
Similar Items
-
Mind the Sim2Real Gap in User Simulation for Agentic Tasks
by: Zhou, Xuhui, et al.
Published: (2026) -
Training Proactive and Personalized LLM Agents
by: Sun, Weiwei, et al.
Published: (2025) -
GradAlign: Gradient-Aligned Data Selection for LLM Reinforcement Learning
by: Yang, Ningyuan, et al.
Published: (2026) -
Ambig-SWE: Interactive Agents to Overcome Underspecificity in Software Engineering
by: Vijayvargiya, Sanidhya, et al.
Published: (2025) -
Optimizing Temperature for Language Models with Multi-Sample Inference
by: Du, Weihua, et al.
Published: (2025)