MARSHAL: Incentivizing Multi-Agent Reasoning via Self-Play with Strategic LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Yuan, Huining, Xu, Zelai, Tan, Zheyue, Yi, Xiangmin, Guang, Mo, Long, Kaiwen, Hui, Haojia, Li, Boxun, Chen, Xinlei, Zhao, Bo, Zhang, Xiao-Ping, Yu, Chao, Wang, Yu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environments
by: Xu, Zelai, et al.
Published: (2025)
by: Xu, Zelai, et al.
Published: (2025)
EARL: Efficient Agentic Reinforcement Learning Systems for Large Language Models
by: Tan, Zheyue, et al.
Published: (2025)
by: Tan, Zheyue, et al.
Published: (2025)
VolleyBots: A Testbed for Multi-Drone Volleyball Game Combining Motion Control and Strategic Play
by: Xu, Zelai, et al.
Published: (2025)
by: Xu, Zelai, et al.
Published: (2025)
Language Agents with Reinforcement Learning for Strategic Play in the Werewolf Game
by: Xu, Zelai, et al.
Published: (2023)
by: Xu, Zelai, et al.
Published: (2023)
Verifiable Process Rewards for Agentic Reasoning
by: Yuan, Huining, et al.
Published: (2026)
by: Yuan, Huining, et al.
Published: (2026)
SATQuest: A Verifier for Logical Reasoning Evaluation and Reinforcement Fine-Tuning of LLMs
by: Zhao, Yanxiao, et al.
Published: (2025)
by: Zhao, Yanxiao, et al.
Published: (2025)
RE-PO: Robust Enhanced Policy Optimization as a General Framework for LLM Alignment
by: Cao, Xiaoyang, et al.
Published: (2025)
by: Cao, Xiaoyang, et al.
Published: (2025)
Learning Strategic Language Agents in the Werewolf Game with Iterative Latent Space Policy Optimization
by: Xu, Zelai, et al.
Published: (2025)
by: Xu, Zelai, et al.
Published: (2025)
SAC Flow: Sample-Efficient Reinforcement Learning of Flow-Based Policies via Velocity-Reparameterized Sequential Modeling
by: Zhang, Yixian, et al.
Published: (2025)
by: Zhang, Yixian, et al.
Published: (2025)
SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning
by: Liu, Bo, et al.
Published: (2025)
by: Liu, Bo, et al.
Published: (2025)
Mastering Multi-Drone Volleyball through Hierarchical Co-Self-Play Reinforcement Learning
by: Zhang, Ruize, et al.
Published: (2025)
by: Zhang, Ruize, et al.
Published: (2025)
Enhancing Language Agent Strategic Reasoning through Self-Play in Adversarial Games
by: Zhang, Yikai, et al.
Published: (2025)
by: Zhang, Yikai, et al.
Published: (2025)
MAGE: Meta-Reinforcement Learning for Language Agents toward Strategic Exploration and Exploitation
by: Yang, Lu, et al.
Published: (2026)
by: Yang, Lu, et al.
Published: (2026)
Incentivizing LLMs to Self-Verify Their Answers
by: Zhang, Fuxiang, et al.
Published: (2025)
by: Zhang, Fuxiang, et al.
Published: (2025)
HiMed: Incentivizing Hindi Reasoning in Medical LLMs
by: Jiang, Dingfeng, et al.
Published: (2026)
by: Jiang, Dingfeng, et al.
Published: (2026)
AED: Automatic Discovery of Effective and Diverse Vulnerabilities for Autonomous Driving Policy with Large Language Models
by: Qiu, Le, et al.
Published: (2025)
by: Qiu, Le, et al.
Published: (2025)
Learning to Lead: Incentivizing Strategic Agents in the Dark
by: Wu, Yuchen, et al.
Published: (2025)
by: Wu, Yuchen, et al.
Published: (2025)
MERRY: Semantically Decoupled Evaluation of Multimodal Emotional and Role Consistencies of Role-Playing Agents
by: Wang, Zhenyu, et al.
Published: (2026)
by: Wang, Zhenyu, et al.
Published: (2026)
AceSearcher: Bootstrapping Reasoning and Search for LLMs via Reinforced Self-Play
by: Xu, Ran, et al.
Published: (2025)
by: Xu, Ran, et al.
Published: (2025)
LearnPruner: Rethinking Attention-based Token Pruning in Vision Language Models
by: Takezoe, Rinyoichi, et al.
Published: (2026)
by: Takezoe, Rinyoichi, et al.
Published: (2026)
Sparks of Cooperative Reasoning: LLMs as Strategic Hanabi Agents
by: Ramesh, Mahesh, et al.
Published: (2026)
by: Ramesh, Mahesh, et al.
Published: (2026)
Strat-Reasoner: Reinforcing Strategic Reasoning of LLMs in Multi-Agent Games
by: He, Yidong, et al.
Published: (2026)
by: He, Yidong, et al.
Published: (2026)
SkillGraph: Self-Evolving Multi-Agent Collaboration with Multimodal Graph Topology
by: Nie, Zheng, et al.
Published: (2026)
by: Nie, Zheng, et al.
Published: (2026)
MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMs
by: Wang, Kevin, et al.
Published: (2026)
by: Wang, Kevin, et al.
Published: (2026)
Crafting Customisable Characters with LLMs: A Persona-Driven Role-Playing Agent Framework
by: Yang, Bohao, et al.
Published: (2024)
by: Yang, Bohao, et al.
Published: (2024)
Trans-Zero: Self-Play Incentivizes Large Language Models for Multilingual Translation Without Parallel Data
by: Zou, Wei, et al.
Published: (2025)
by: Zou, Wei, et al.
Published: (2025)
Cogito, Ergo Ludo: An Agent that Learns to Play by Reasoning and Planning
by: Wang, Sai, et al.
Published: (2025)
by: Wang, Sai, et al.
Published: (2025)
The Phylogeography of Selfing Caulokaempferia coenobialis Responded to Pleistocene Karst Development and Transgressions‐Regressions in Southern China
by: Guo‐Hui Lu, et al.
Published: (2025)
by: Guo‐Hui Lu, et al.
Published: (2025)
Vision-Zero: Scalable VLM Self-Improvement via Strategic Gamified Self-Play
by: Wang, Qinsi, et al.
Published: (2025)
by: Wang, Qinsi, et al.
Published: (2025)
LLM-Gomoku: A Large Language Model-Based System for Strategic Gomoku with Self-Play and Reinforcement Learning
by: Wang, Hui
Published: (2025)
by: Wang, Hui
Published: (2025)
Two Tales of Persona in LLMs: A Survey of Role-Playing and Personalization
by: Tseng, Yu-Min, et al.
Published: (2024)
by: Tseng, Yu-Min, et al.
Published: (2024)
Incentives and Strategic Behaviour: An Experiment
by: Esteban-Casanelles, Teresa, et al.
Published: (2023)
by: Esteban-Casanelles, Teresa, et al.
Published: (2023)
Anticipating Gaming to Incentivize Improvement: Guiding Agents in (Fair) Strategic Classification
by: Alhanouti, Sura, et al.
Published: (2025)
by: Alhanouti, Sura, et al.
Published: (2025)
Robust and Performance Incentivizing Algorithms for Multi-Armed Bandits with Strategic Agents
by: Esmaeili, Seyed A., et al.
Published: (2023)
by: Esmaeili, Seyed A., et al.
Published: (2023)
Hypergame Rationalisability: Solving Agent Misalignment In Strategic Play
by: Trencsenyi, Vince
Published: (2025)
by: Trencsenyi, Vince
Published: (2025)
Playing Language Game with LLMs Leads to Jailbreaking
by: Peng, Yu, et al.
Published: (2024)
by: Peng, Yu, et al.
Published: (2024)
ReasonCache: Accelerating Large Reasoning Model Serving through KV Cache Sharing
by: Chen, Kaiwen, et al.
Published: (2025)
by: Chen, Kaiwen, et al.
Published: (2025)
WideSeek-R1: Exploring Width Scaling for Broad Information Seeking via Multi-Agent Reinforcement Learning
by: Xu, Zelai, et al.
Published: (2026)
by: Xu, Zelai, et al.
Published: (2026)
LV-Eval: A Balanced Long-Context Benchmark with 5 Length Levels Up to 256K
by: Yuan, Tao, et al.
Published: (2024)
by: Yuan, Tao, et al.
Published: (2024)
RAG-R1: Incentivizing the Search and Reasoning Capabilities of LLMs through Multi-query Parallelism
by: Tan, Zhiwen, et al.
Published: (2025)
by: Tan, Zhiwen, et al.
Published: (2025)
Similar Items
-
VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environments
by: Xu, Zelai, et al.
Published: (2025) -
EARL: Efficient Agentic Reinforcement Learning Systems for Large Language Models
by: Tan, Zheyue, et al.
Published: (2025) -
VolleyBots: A Testbed for Multi-Drone Volleyball Game Combining Motion Control and Strategic Play
by: Xu, Zelai, et al.
Published: (2025) -
Language Agents with Reinforcement Learning for Strategic Play in the Werewolf Game
by: Xu, Zelai, et al.
Published: (2023) -
Verifiable Process Rewards for Agentic Reasoning
by: Yuan, Huining, et al.
Published: (2026)