EgoSocialArena: Benchmarking the Social Intelligence of Large Language Models from a First-person Perspective
Fuente:
arXiv
Saved in:
| Main Authors: | Hou, Guiyang, Zhang, Wenqi, Shen, Yongliang, Tan, Zeqi, Shen, Sihao, Lu, Weiming |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Active Confusion Expression in Large Language Models: Leveraging World Models toward Better Social Reasoning
by: Du, Jialu, et al.
Published: (2025)
by: Du, Jialu, et al.
Published: (2025)
Data-Copilot: Bridging Billions of Data and Humans with Autonomous Workflow
by: Zhang, Wenqi, et al.
Published: (2023)
by: Zhang, Wenqi, et al.
Published: (2023)
Cooper: Co-Optimizing Policy and Reward Models in Reinforcement Learning for Large Language Models
by: Hong, Haitao, et al.
Published: (2025)
by: Hong, Haitao, et al.
Published: (2025)
Agent-Pro: Learning to Evolve via Policy-Level Reflection and Optimization
by: Zhang, Wenqi, et al.
Published: (2024)
by: Zhang, Wenqi, et al.
Published: (2024)
TimeHC-RL: Temporal-aware Hierarchical Cognitive Reinforcement Learning for Enhancing LLMs' Social Intelligence
by: Hou, Guiyang, et al.
Published: (2025)
by: Hou, Guiyang, et al.
Published: (2025)
TaskBench: Benchmarking Large Language Models for Task Automation
by: Shen, Yongliang, et al.
Published: (2023)
by: Shen, Yongliang, et al.
Published: (2023)
TimeToM: Temporal Space is the Key to Unlocking the Door of Large Language Models' Theory-of-Mind
by: Hou, Guiyang, et al.
Published: (2024)
by: Hou, Guiyang, et al.
Published: (2024)
Advancing Process Verification for Large Language Models via Tree-Based Preference Learning
by: He, Mingqian, et al.
Published: (2024)
by: He, Mingqian, et al.
Published: (2024)
CoVerRL: Breaking the Consensus Trap in Label-Free Reasoning via Generator-Verifier Co-Evolution
by: Pan, Teng, et al.
Published: (2026)
by: Pan, Teng, et al.
Published: (2026)
Mind the Gap: Bridging Thought Leap for Improved Chain-of-Thought Tuning
by: Xu, Haolei, et al.
Published: (2025)
by: Xu, Haolei, et al.
Published: (2025)
Self-Contrast: Better Reflection Through Inconsistent Solving Perspectives
by: Zhang, Wenqi, et al.
Published: (2024)
by: Zhang, Wenqi, et al.
Published: (2024)
ViewSpatial-Bench: Evaluating Multi-perspective Spatial Localization in Vision-Language Models
by: Li, Dingming, et al.
Published: (2025)
by: Li, Dingming, et al.
Published: (2025)
Think Twice, Click Once: Enhancing GUI Grounding via Fast and Slow Systems
by: Tang, Fei, et al.
Published: (2025)
by: Tang, Fei, et al.
Published: (2025)
A Survey on (M)LLM-Based GUI Agents
by: Tang, Fei, et al.
Published: (2025)
by: Tang, Fei, et al.
Published: (2025)
EasySteer: A Unified Framework for High-Performance and Extensible LLM Steering
by: Xu, Haolei, et al.
Published: (2025)
by: Xu, Haolei, et al.
Published: (2025)
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models
by: Zhu, Yakun, et al.
Published: (2025)
by: Zhu, Yakun, et al.
Published: (2025)
AskToAct: Enhancing LLMs Tool Use via Self-Correcting Clarification
by: Zhang, Xuan, et al.
Published: (2025)
by: Zhang, Xuan, et al.
Published: (2025)
SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language Models
by: Li, Hongxing, et al.
Published: (2025)
by: Li, Hongxing, et al.
Published: (2025)
Measuring Social Norms of Large Language Models
by: Yuan, Ye, et al.
Published: (2024)
by: Yuan, Ye, et al.
Published: (2024)
LongReasonArena: A Long Reasoning Benchmark for Large Language Models
by: Ding, Jiayu, et al.
Published: (2025)
by: Ding, Jiayu, et al.
Published: (2025)
Let LRMs Break Free from Overthinking via Self-Braking Tuning
by: Zhao, Haoran, et al.
Published: (2025)
by: Zhao, Haoran, et al.
Published: (2025)
VerifyBench: Benchmarking Reference-based Reward Systems for Large Language Models
by: Yan, Yuchen, et al.
Published: (2025)
by: Yan, Yuchen, et al.
Published: (2025)
GaVaMoE: Gaussian-Variational Gated Mixture of Experts for Explainable Recommendation
by: Tang, Fei, et al.
Published: (2024)
by: Tang, Fei, et al.
Published: (2024)
EgoNormia: Benchmarking Physical Social Norm Understanding
by: Rezaei, MohammadHossein, et al.
Published: (2025)
by: Rezaei, MohammadHossein, et al.
Published: (2025)
Towards Objectively Benchmarking Social Intelligence for Language Agents at Action Level
by: Wang, Chenxu, et al.
Published: (2024)
by: Wang, Chenxu, et al.
Published: (2024)
InftyThink: Breaking the Length Limits of Long-Context Reasoning in Large Language Models
by: Yan, Yuchen, et al.
Published: (2025)
by: Yan, Yuchen, et al.
Published: (2025)
GroundAct: Can LLM Agents Ground Actions in Environmental States?
by: Wang, Zixuan, et al.
Published: (2025)
by: Wang, Zixuan, et al.
Published: (2025)
Milestone-Guided Policy Learning for Long-Horizon Language Agents
by: Wang, Zixuan, et al.
Published: (2026)
by: Wang, Zixuan, et al.
Published: (2026)
Time Is a Feature: Exploiting Temporal Dynamics in Diffusion Language Models
by: Wang, Wen, et al.
Published: (2025)
by: Wang, Wen, et al.
Published: (2025)
Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena
by: Myrzakhan, Aidar, et al.
Published: (2024)
by: Myrzakhan, Aidar, et al.
Published: (2024)
PokeLLMon: A Human-Parity Agent for Pokemon Battles with Large Language Models
by: Hu, Sihao, et al.
Published: (2024)
by: Hu, Sihao, et al.
Published: (2024)
SVGenius: Benchmarking LLMs in SVG Understanding, Editing and Generation
by: Chen, Siqi, et al.
Published: (2025)
by: Chen, Siqi, et al.
Published: (2025)
GSM8K-V: Can Vision Language Models Solve Grade School Math Word Problems in Visual Contexts
by: Yuan, Fan, et al.
Published: (2025)
by: Yuan, Fan, et al.
Published: (2025)
ResearchArena: Benchmarking Large Language Models' Ability to Collect and Organize Information as Research Agents
by: Kang, Hao, et al.
Published: (2024)
by: Kang, Hao, et al.
Published: (2024)
A Perspective on Large Language Models, Intelligent Machines, and Knowledge Acquisition
by: Cherkassky, Vladimir, et al.
Published: (2024)
by: Cherkassky, Vladimir, et al.
Published: (2024)
GraphArena: Evaluating and Exploring Large Language Models on Graph Computation
by: Tang, Jianheng, et al.
Published: (2024)
by: Tang, Jianheng, et al.
Published: (2024)
Do LLMs Understand Social Knowledge? Evaluating the Sociability of Large Language Models with SocKET Benchmark
by: Choi, Minje, et al.
Published: (2023)
by: Choi, Minje, et al.
Published: (2023)
Benchmarking Large Language Models on Reference Extraction and Parsing in the Social Sciences and Humanities
by: Zhu, Yurui, et al.
Published: (2026)
by: Zhu, Yurui, et al.
Published: (2026)
SocialVeil: Probing Social Intelligence of Language Agents under Communication Barriers
by: Xuan, Keyang, et al.
Published: (2026)
by: Xuan, Keyang, et al.
Published: (2026)
Mobile-MMLU: A Mobile Intelligence Language Understanding Benchmark
by: Bsharat, Sondos Mahmoud, et al.
Published: (2025)
by: Bsharat, Sondos Mahmoud, et al.
Published: (2025)
Similar Items
-
Active Confusion Expression in Large Language Models: Leveraging World Models toward Better Social Reasoning
by: Du, Jialu, et al.
Published: (2025) -
Data-Copilot: Bridging Billions of Data and Humans with Autonomous Workflow
by: Zhang, Wenqi, et al.
Published: (2023) -
Cooper: Co-Optimizing Policy and Reward Models in Reinforcement Learning for Large Language Models
by: Hong, Haitao, et al.
Published: (2025) -
Agent-Pro: Learning to Evolve via Policy-Level Reflection and Optimization
by: Zhang, Wenqi, et al.
Published: (2024) -
TimeHC-RL: Temporal-aware Hierarchical Cognitive Reinforcement Learning for Enhancing LLMs' Social Intelligence
by: Hou, Guiyang, et al.
Published: (2025)