M3-BENCH: Process-Aware Evaluation of LLM Agents' Social Behaviors in Mixed-Motive Games
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Xie, Sixiong, Shi, Zhuofan, Shen, Haiyang, Ma, Yun, Jing, Xiang |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
SGR-Bench: Benchmarking Search Agents on State-Gated Retrieval
par: Li, Ningyuan, et autres
Publié: (2026)
par: Li, Ningyuan, et autres
Publié: (2026)
DeepWeb-Bench: A Deep Research Benchmark Demanding Massive Cross-Source Evidence and Long-Horizon Derivation
par: Xie, Sixiong, et autres
Publié: (2026)
par: Xie, Sixiong, et autres
Publié: (2026)
MindLoom: Composing Thought Modes for Frontier-Level Reasoning Data Synthesis
par: Shen, Haiyang, et autres
Publié: (2026)
par: Shen, Haiyang, et autres
Publié: (2026)
LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research
par: Yan, Shuo, et autres
Publié: (2025)
par: Yan, Shuo, et autres
Publié: (2025)
Explaining Decisions of Agents in Mixed-Motive Games
par: Orner, Maayan, et autres
Publié: (2024)
par: Orner, Maayan, et autres
Publié: (2024)
ViDR: Grounding Multimodal Deep Research Reports in Source Visual Evidence
par: Shi, Zhuofan, et autres
Publié: (2026)
par: Shi, Zhuofan, et autres
Publié: (2026)
OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces
par: Li, Xiaozhe, et autres
Publié: (2026)
par: Li, Xiaozhe, et autres
Publié: (2026)
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems
par: Li, Xiaozhe, et autres
Publié: (2025)
par: Li, Xiaozhe, et autres
Publié: (2025)
PAC-BENCH: Evaluating Multi-Agent Collaboration under Privacy Constraints
par: Park, Minjun, et autres
Publié: (2026)
par: Park, Minjun, et autres
Publié: (2026)
DVM: Towards Controllable LLM Agents in Social Deduction Games
par: Zhang, Zheng, et autres
Publié: (2025)
par: Zhang, Zheng, et autres
Publié: (2025)
Minding Motivation: The Effect of Intrinsic Motivation on Agent Behaviors
par: Villalobos-Arias, Leonardo, et autres
Publié: (2025)
par: Villalobos-Arias, Leonardo, et autres
Publié: (2025)
Evaluating Generalization Capabilities of LLM-Based Agents in Mixed-Motive Scenarios Using Concordia
par: Smith, Chandler, et autres
Publié: (2025)
par: Smith, Chandler, et autres
Publié: (2025)
Adaptive Punishment for Cooperation in Mixed-Motive Games
par: Tang, Min, et autres
Publié: (2026)
par: Tang, Min, et autres
Publié: (2026)
ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning
par: Potamitis, Nearchos, et autres
Publié: (2025)
par: Potamitis, Nearchos, et autres
Publié: (2025)
Constrained Intrinsic Motivation for Reinforcement Learning
par: Zheng, Xiang, et autres
Publié: (2024)
par: Zheng, Xiang, et autres
Publié: (2024)
FAIRGAMER: Evaluating Social Biases in LLM-Based Video Game NPCs
par: Shi, Bingkang, et autres
Publié: (2025)
par: Shi, Bingkang, et autres
Publié: (2025)
Can Machines Think Like Humans? A Behavioral Evaluation of LLM Agents in Dictator Games
par: Ma, Ji
Publié: (2024)
par: Ma, Ji
Publié: (2024)
Learning to Balance Altruism and Self-interest Based on Empathy in Mixed-Motive Games
par: Kong, Fanqi, et autres
Publié: (2024)
par: Kong, Fanqi, et autres
Publié: (2024)
Benevolent Dictators? On LLM Agent Behavior in Dictator Games
par: Einwiller, Andreas, et autres
Publié: (2025)
par: Einwiller, Andreas, et autres
Publié: (2025)
GVGAI-LLM: Evaluating Large Language Model Agents with Infinite Games
par: Li, Yuchen, et autres
Publié: (2025)
par: Li, Yuchen, et autres
Publié: (2025)
How Memory Management Impacts LLM Agents: An Empirical Study of Experience-Following Behavior
par: Xiong, Zidi, et autres
Publié: (2025)
par: Xiong, Zidi, et autres
Publié: (2025)
ReplicatorBench: Benchmarking LLM Agents for Replicability in Social and Behavioral Sciences
par: Nguyen, Bang, et autres
Publié: (2026)
par: Nguyen, Bang, et autres
Publié: (2026)
GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents
par: Costarelli, Anthony, et autres
Publié: (2024)
par: Costarelli, Anthony, et autres
Publié: (2024)
GameArena: Evaluating LLM Reasoning through Live Computer Games
par: Hu, Lanxiang, et autres
Publié: (2024)
par: Hu, Lanxiang, et autres
Publié: (2024)
CreativeGame:Toward Mechanic-Aware Creative Game Generation
par: Ma, Hongnan, et autres
Publié: (2026)
par: Ma, Hongnan, et autres
Publié: (2026)
M-$LLM^3$REC: A Motivation-Aware User-Item Interaction Framework for Enhancing Recommendation Accuracy with LLMs
par: Chen, Lining, et autres
Publié: (2025)
par: Chen, Lining, et autres
Publié: (2025)
Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games
par: Park, Dongmin, et autres
Publié: (2025)
par: Park, Dongmin, et autres
Publié: (2025)
FLOW-BENCH: Towards Conversational Generation of Enterprise Workflows
par: Duesterwald, Evelyn, et autres
Publié: (2025)
par: Duesterwald, Evelyn, et autres
Publié: (2025)
Game Agent Driven by Free-Form Text Command: Using LLM-based Code Generation and Behavior Branch
par: Ito, Ray, et autres
Publié: (2024)
par: Ito, Ray, et autres
Publié: (2024)
ClawTrace: Cost-Aware Tracing for LLM Agent Skill Distillation
par: Yuan, Boqin, et autres
Publié: (2026)
par: Yuan, Boqin, et autres
Publié: (2026)
JudgeAgent: Beyond Static Benchmarks for Knowledge-Driven and Dynamic LLM Evaluation
par: Shi, Zhichao, et autres
Publié: (2025)
par: Shi, Zhichao, et autres
Publié: (2025)
Deciphering Digital Detectives: Understanding LLM Behaviors and Capabilities in Multi-Agent Mystery Games
par: Wu, Dekun, et autres
Publié: (2023)
par: Wu, Dekun, et autres
Publié: (2023)
TRACE: Trajectory Risk-Aware Compression for Long-Horizon Agent Safety
par: Hong, Zhepei, et autres
Publié: (2026)
par: Hong, Zhepei, et autres
Publié: (2026)
Do LLM Agents Exhibit Social Behavior?
par: Leng, Yan, et autres
Publié: (2023)
par: Leng, Yan, et autres
Publié: (2023)
Evaluate-as-Action: Self-Evaluated Process Rewards for Retrieval-Augmented Agents
par: Shu, Jiangming, et autres
Publié: (2026)
par: Shu, Jiangming, et autres
Publié: (2026)
RUST-BENCH: Benchmarking LLM Reasoning on Unstructured Text within Structured Tables
par: Abhyankar, Nikhil, et autres
Publié: (2025)
par: Abhyankar, Nikhil, et autres
Publié: (2025)
LifelongAgentBench: Evaluating LLM Agents as Lifelong Learners
par: Zheng, Junhao, et autres
Publié: (2025)
par: Zheng, Junhao, et autres
Publié: (2025)
The Behavior Gap: Evaluating Zero-shot LLM Agents in Complex Task-Oriented Dialogs
par: Baidya, Avinash, et autres
Publié: (2025)
par: Baidya, Avinash, et autres
Publié: (2025)
Experience Transfer for Multimodal LLM Agents in Minecraft Game
par: Li, Chenghao, et autres
Publié: (2026)
par: Li, Chenghao, et autres
Publié: (2026)
AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents
par: Feng, Yunhao, et autres
Publié: (2026)
par: Feng, Yunhao, et autres
Publié: (2026)
Documents similaires
-
SGR-Bench: Benchmarking Search Agents on State-Gated Retrieval
par: Li, Ningyuan, et autres
Publié: (2026) -
DeepWeb-Bench: A Deep Research Benchmark Demanding Massive Cross-Source Evidence and Long-Horizon Derivation
par: Xie, Sixiong, et autres
Publié: (2026) -
MindLoom: Composing Thought Modes for Frontier-Level Reasoning Data Synthesis
par: Shen, Haiyang, et autres
Publié: (2026) -
LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research
par: Yan, Shuo, et autres
Publié: (2025) -
Explaining Decisions of Agents in Mixed-Motive Games
par: Orner, Maayan, et autres
Publié: (2024)