SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Xuhui, Zhu, Hao, Mathur, Leena, Zhang, Ruohong, Yu, Haofei, Qi, Zhengyang, Morency, Louis-Philippe, Bisk, Yonatan, Fried, Daniel, Neubig, Graham, Sap, Maarten |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SOTOPIA-$π$: Interactive Learning of Socially Intelligent Language Agents
von: Wang, Ruiyi, et al.
Veröffentlicht: (2024)
von: Wang, Ruiyi, et al.
Veröffentlicht: (2024)
When Should AI Read the Room? Public Perceptions of Social Intelligence in AI Agents
von: Mathur, Leena, et al.
Veröffentlicht: (2026)
von: Mathur, Leena, et al.
Veröffentlicht: (2026)
Advancing Social Intelligence in AI Agents: Technical Challenges and Open Questions
von: Mathur, Leena, et al.
Veröffentlicht: (2024)
von: Mathur, Leena, et al.
Veröffentlicht: (2024)
SOTOPIA-TOM: Evaluating Information Management in Multi-Agent Interaction with Theory of Mind
von: YS, Yashwanth, et al.
Veröffentlicht: (2026)
von: YS, Yashwanth, et al.
Veröffentlicht: (2026)
Ambig-SWE: Interactive Agents to Overcome Underspecificity in Software Engineering
von: Vijayvargiya, Sanidhya, et al.
Veröffentlicht: (2025)
von: Vijayvargiya, Sanidhya, et al.
Veröffentlicht: (2025)
Social Caption: Evaluating Social Understanding in Multimodal Models
von: Thumu, Bhaavanaa, et al.
Veröffentlicht: (2026)
von: Thumu, Bhaavanaa, et al.
Veröffentlicht: (2026)
LIFELONG SOTOPIA: Evaluating Social Intelligence of Language Agents Over Lifelong Social Interactions
von: Goel, Hitesh, et al.
Veröffentlicht: (2025)
von: Goel, Hitesh, et al.
Veröffentlicht: (2025)
Social Genome: Grounded Social Reasoning Abilities of Multimodal Models
von: Mathur, Leena, et al.
Veröffentlicht: (2025)
von: Mathur, Leena, et al.
Veröffentlicht: (2025)
HEMM: Holistic Evaluation of Multimodal Foundation Models
von: Liang, Paul Pu, et al.
Veröffentlicht: (2024)
von: Liang, Paul Pu, et al.
Veröffentlicht: (2024)
MMoE: Enhancing Multimodal Models with Mixtures of Multimodal Interaction Experts
von: Yu, Haofei, et al.
Veröffentlicht: (2023)
von: Yu, Haofei, et al.
Veröffentlicht: (2023)
TOM-SWE: User Mental Modeling For Software Engineering Agents
von: Zhou, Xuhui, et al.
Veröffentlicht: (2025)
von: Zhou, Xuhui, et al.
Veröffentlicht: (2025)
OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety
von: Vijayvargiya, Sanidhya, et al.
Veröffentlicht: (2025)
von: Vijayvargiya, Sanidhya, et al.
Veröffentlicht: (2025)
OpenFace 3.0: A Lightweight Multitask System for Comprehensive Facial Behavior Analysis
von: Hu, Jiewen, et al.
Veröffentlicht: (2025)
von: Hu, Jiewen, et al.
Veröffentlicht: (2025)
SoMi-ToM: Evaluating Multi-Perspective Theory of Mind in Embodied Social Interactions
von: Fan, Xianzhe, et al.
Veröffentlicht: (2025)
von: Fan, Xianzhe, et al.
Veröffentlicht: (2025)
From Reproduction to Replication: Evaluating Research Agents with Progressive Code Masking
von: Kim, Gyeongwon James, et al.
Veröffentlicht: (2025)
von: Kim, Gyeongwon James, et al.
Veröffentlicht: (2025)
WebArena: A Realistic Web Environment for Building Autonomous Agents
von: Zhou, Shuyan, et al.
Veröffentlicht: (2023)
von: Zhou, Shuyan, et al.
Veröffentlicht: (2023)
Training Proactive and Personalized LLM Agents
von: Sun, Weiwei, et al.
Veröffentlicht: (2025)
von: Sun, Weiwei, et al.
Veröffentlicht: (2025)
SOTOPIA-S4: a user-friendly system for flexible, customizable, and large-scale social simulation
von: Zhou, Xuhui, et al.
Veröffentlicht: (2025)
von: Zhou, Xuhui, et al.
Veröffentlicht: (2025)
Is this the real life? Is this just fantasy? The Misleading Success of Simulating Social Interactions With LLMs
von: Zhou, Xuhui, et al.
Veröffentlicht: (2024)
von: Zhou, Xuhui, et al.
Veröffentlicht: (2024)
Is the Pope Catholic? Yes, the Pope is Catholic. Generative Evaluation of Non-Literal Intent Resolution in LLMs
von: Yerukola, Akhila, et al.
Veröffentlicht: (2024)
von: Yerukola, Akhila, et al.
Veröffentlicht: (2024)
Stereotype or Personalization? User Identity Biases Chatbot Recommendations
von: Kantharuban, Anjali, et al.
Veröffentlicht: (2024)
von: Kantharuban, Anjali, et al.
Veröffentlicht: (2024)
AutoPresent: Designing Structured Visuals from Scratch
von: Ge, Jiaxin, et al.
Veröffentlicht: (2025)
von: Ge, Jiaxin, et al.
Veröffentlicht: (2025)
Social World Models
von: Zhou, Xuhui, et al.
Veröffentlicht: (2025)
von: Zhou, Xuhui, et al.
Veröffentlicht: (2025)
Agent Workflow Memory
von: Wang, Zora Zhiruo, et al.
Veröffentlicht: (2024)
von: Wang, Zora Zhiruo, et al.
Veröffentlicht: (2024)
SOTOPIA-$Ω$: Dynamic Strategy Injection Learning and Social Instruction Following Evaluation for Social Agents
von: Zhang, Wenyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Wenyuan, et al.
Veröffentlicht: (2025)
TroVE: Inducing Verifiable and Efficient Toolboxes for Solving Programmatic Tasks
von: Wang, Zhiruo, et al.
Veröffentlicht: (2024)
von: Wang, Zhiruo, et al.
Veröffentlicht: (2024)
1-2-3 Check: Enhancing Contextual Privacy in LLM via Multi-Agent Reasoning
von: Li, Wenkai, et al.
Veröffentlicht: (2025)
von: Li, Wenkai, et al.
Veröffentlicht: (2025)
What Are Tools Anyway? A Survey from the Language Model Perspective
von: Wang, Zhiruo, et al.
Veröffentlicht: (2024)
von: Wang, Zhiruo, et al.
Veröffentlicht: (2024)
Learning Model Successors
von: Chang, Yingshan, et al.
Veröffentlicht: (2025)
von: Chang, Yingshan, et al.
Veröffentlicht: (2025)
Language Models Need Inductive Biases to Count Inductively
von: Chang, Yingshan, et al.
Veröffentlicht: (2024)
von: Chang, Yingshan, et al.
Veröffentlicht: (2024)
Communicate-Predict-Act: Evaluating Social Intelligence of Agents
von: Shoresh, David, et al.
Veröffentlicht: (2026)
von: Shoresh, David, et al.
Veröffentlicht: (2026)
Unsupervised Discovery of Long-Term Spatiotemporal Periodic Workflows in Human Activities
von: Yang, Fan, et al.
Veröffentlicht: (2025)
von: Yang, Fan, et al.
Veröffentlicht: (2025)
REM: Evaluating LLM Embodied Spatial Reasoning through Multi-Frame Trajectories
von: Thompson, Jacob, et al.
Veröffentlicht: (2025)
von: Thompson, Jacob, et al.
Veröffentlicht: (2025)
Sotopia-RL: Reward Design for Social Intelligence
von: Yu, Haofei, et al.
Veröffentlicht: (2025)
von: Yu, Haofei, et al.
Veröffentlicht: (2025)
PolygloToxicityPrompts: Multilingual Evaluation of Neural Toxic Degeneration in Large Language Models
von: Jain, Devansh, et al.
Veröffentlicht: (2024)
von: Jain, Devansh, et al.
Veröffentlicht: (2024)
Training Versatile Coding Agents in Synthetic Environments
von: Zhu, Yiqi, et al.
Veröffentlicht: (2025)
von: Zhu, Yiqi, et al.
Veröffentlicht: (2025)
Examining the Effect of Explanations of AI Privacy Redaction in AI-mediated Interactions
von: Kaushik, Roshni, et al.
Veröffentlicht: (2026)
von: Kaushik, Roshni, et al.
Veröffentlicht: (2026)
Rethinking Theory of Mind Benchmarks for LLMs: Towards A User-Centered Perspective
von: Wang, Qiaosi, et al.
Veröffentlicht: (2025)
von: Wang, Qiaosi, et al.
Veröffentlicht: (2025)
GoodPoint: Learning Constructive Scientific Paper Feedback from Author Responses
von: Mun, Jimin, et al.
Veröffentlicht: (2026)
von: Mun, Jimin, et al.
Veröffentlicht: (2026)
ANAVI: Audio Noise Awareness using Visuals of Indoor environments for NAVIgation
von: Jain, Vidhi, et al.
Veröffentlicht: (2024)
von: Jain, Vidhi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
SOTOPIA-$π$: Interactive Learning of Socially Intelligent Language Agents
von: Wang, Ruiyi, et al.
Veröffentlicht: (2024) -
When Should AI Read the Room? Public Perceptions of Social Intelligence in AI Agents
von: Mathur, Leena, et al.
Veröffentlicht: (2026) -
Advancing Social Intelligence in AI Agents: Technical Challenges and Open Questions
von: Mathur, Leena, et al.
Veröffentlicht: (2024) -
SOTOPIA-TOM: Evaluating Information Management in Multi-Agent Interaction with Theory of Mind
von: YS, Yashwanth, et al.
Veröffentlicht: (2026) -
Ambig-SWE: Interactive Agents to Overcome Underspecificity in Software Engineering
von: Vijayvargiya, Sanidhya, et al.
Veröffentlicht: (2025)