FM3Q: Factorized Multi-Agent MiniMax Q-Learning for Two-Team Zero-Sum Markov Game
Fuente:
arXiv
Guardado en:
| Autores principales: | Hu, Guangzheng, Zhu, Yuanheng, Li, Haoran, Zhao, Dongbin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence
por: MiniMax, et al.
Publicado: (2026)
por: MiniMax, et al.
Publicado: (2026)
Discretizing Continuous Action Space with Unimodal Probability Distributions for On-Policy Reinforcement Learning
por: Zhu, Yuanyang, et al.
Publicado: (2024)
por: Zhu, Yuanyang, et al.
Publicado: (2024)
A Hierarchical Deep Reinforcement Learning Framework for 6-DOF UCAV Air-to-Air Combat
por: Chai, Jiajun, et al.
Publicado: (2022)
por: Chai, Jiajun, et al.
Publicado: (2022)
Computing Ex Ante Equilibrium in Heterogeneous Zero-Sum Team Games
por: Liu, Naming, et al.
Publicado: (2024)
por: Liu, Naming, et al.
Publicado: (2024)
Bilevel Optimization over Saddle Points of Zero-Sum Markov Games
por: Zheng, Zihao, et al.
Publicado: (2026)
por: Zheng, Zihao, et al.
Publicado: (2026)
RLAE: Reinforcement Learning-Assisted Ensemble for LLMs
por: Fu, Yuqian, et al.
Publicado: (2025)
por: Fu, Yuqian, et al.
Publicado: (2025)
MinMaxMin $Q$-learning
por: Soffair, Nitsan, et al.
Publicado: (2024)
por: Soffair, Nitsan, et al.
Publicado: (2024)
High-order Interactions Modeling for Interpretable Multi-Agent Q-Learning
por: Xu, Qinyu, et al.
Publicado: (2025)
por: Xu, Qinyu, et al.
Publicado: (2025)
Global Policy-Space Response Oracles for Two-Player Zero-Sum Games
por: Zhang, Junyu, et al.
Publicado: (2026)
por: Zhang, Junyu, et al.
Publicado: (2026)
DipLLM: Fine-Tuning LLM for Strategic Decision-making in Diplomacy
por: Xu, Kaixuan, et al.
Publicado: (2025)
por: Xu, Kaixuan, et al.
Publicado: (2025)
Mitigating Backdoor Attacks in Federated Learning Using PPA and MiniMax Game Theory
por: Wehbi, Osama, et al.
Publicado: (2026)
por: Wehbi, Osama, et al.
Publicado: (2026)
RiskQ: Risk-sensitive Multi-Agent Reinforcement Learning Value Factorization
por: Shen, Siqi, et al.
Publicado: (2023)
por: Shen, Siqi, et al.
Publicado: (2023)
Zero-Sum Positional Differential Games as a Framework for Robust Reinforcement Learning: Deep Q-Learning Approach
por: Plaksin, Anton, et al.
Publicado: (2024)
por: Plaksin, Anton, et al.
Publicado: (2024)
The FM Agent
por: Li, Annan, et al.
Publicado: (2025)
por: Li, Annan, et al.
Publicado: (2025)
SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning
por: Liu, Bo, et al.
Publicado: (2025)
por: Liu, Bo, et al.
Publicado: (2025)
Improving Sample Efficiency of Model-Free Algorithms for Zero-Sum Markov Games
por: Feng, Songtao, et al.
Publicado: (2023)
por: Feng, Songtao, et al.
Publicado: (2023)
MemQ: Integrating Q-Learning into Self-Evolving Memory Agents over Provenance DAGs
por: Liao, Junwei, et al.
Publicado: (2026)
por: Liao, Junwei, et al.
Publicado: (2026)
ARAC: Adaptive Regularized Multi-Agent Soft Actor-Critic in Graph-Structured Adversarial Games
por: Shi, Ruochuan, et al.
Publicado: (2025)
por: Shi, Ruochuan, et al.
Publicado: (2025)
MiniMax-01: Scaling Foundation Models with Lightning Attention
por: MiniMax, et al.
Publicado: (2025)
por: MiniMax, et al.
Publicado: (2025)
Multi-Agent Inverse Q-Learning from Demonstrations
por: Haynam, Nathaniel, et al.
Publicado: (2025)
por: Haynam, Nathaniel, et al.
Publicado: (2025)
A Multi-Step Minimax Q-learning Algorithm for Two-Player Zero-Sum Markov Games
por: R, Shreyas S, et al.
Publicado: (2024)
por: R, Shreyas S, et al.
Publicado: (2024)
M3GCLR: Multi-View Mini-Max Infinite Skeleton-Data Game Contrastive Learning For Skeleton-Based Action Recognition
por: Li, Yanshan, et al.
Publicado: (2026)
por: Li, Yanshan, et al.
Publicado: (2026)
MiniZero: Comparative Analysis of AlphaZero and MuZero on Go, Othello, and Atari Games
por: Wu, Ti-Rong, et al.
Publicado: (2023)
por: Wu, Ti-Rong, et al.
Publicado: (2023)
Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes
por: Fu, Yuqian, et al.
Publicado: (2026)
por: Fu, Yuqian, et al.
Publicado: (2026)
Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents
por: Putta, Pranav, et al.
Publicado: (2024)
por: Putta, Pranav, et al.
Publicado: (2024)
Q-ARDNS-Multi: A Multi-Agent Quantum Reinforcement Learning Framework with Meta-Cognitive Adaptation for Complex 3D Environments
por: de Sousa, Umberto Gonçalves
Publicado: (2025)
por: de Sousa, Umberto Gonçalves
Publicado: (2025)
Embodiment-Induced Coordination Regimes in Tabular Multi-Agent Q-Learning
por: Atif, Muhammad Ahmed, et al.
Publicado: (2026)
por: Atif, Muhammad Ahmed, et al.
Publicado: (2026)
Learning To Play Atari Games Using Dueling Q-Learning and Hebbian Plasticity
por: Salehin, Md Ashfaq
Publicado: (2024)
por: Salehin, Md Ashfaq
Publicado: (2024)
TeViR: Text-to-Video Reward with Diffusion Models for Efficient Reinforcement Learning
por: Chen, Yuhui, et al.
Publicado: (2025)
por: Chen, Yuhui, et al.
Publicado: (2025)
$\widetilde{O}(T^{-1})$ Convergence to (Coarse) Correlated Equilibria in Full-Information General-Sum Markov Games
por: Mao, Weichao, et al.
Publicado: (2024)
por: Mao, Weichao, et al.
Publicado: (2024)
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
por: Zhang, Bowen, et al.
Publicado: (2025)
por: Zhang, Bowen, et al.
Publicado: (2025)
Cross-domain Random Pre-training with Prototypes for Reinforcement Learning
por: Liu, Xin, et al.
Publicado: (2023)
por: Liu, Xin, et al.
Publicado: (2023)
COvolve: Adversarial Co-Evolution of Large-Language-Model-Generated Policies and Environments via Two-Player Zero-Sum Game
por: Sygkounas, Alkis, et al.
Publicado: (2026)
por: Sygkounas, Alkis, et al.
Publicado: (2026)
MiniMax Entropy Network: Learning Category-Invariant Features for Domain Adaptation
por: Tao, Chaofan, et al.
Publicado: (2019)
por: Tao, Chaofan, et al.
Publicado: (2019)
Drift Q-Learning
por: Houssaini, Anas, et al.
Publicado: (2026)
por: Houssaini, Anas, et al.
Publicado: (2026)
Frictional Q-Learning
por: Kim, Hyunwoo, et al.
Publicado: (2025)
por: Kim, Hyunwoo, et al.
Publicado: (2025)
Flow Q-Learning
por: Park, Seohong, et al.
Publicado: (2025)
por: Park, Seohong, et al.
Publicado: (2025)
MiniMax-Remover: Taming Bad Noise Helps Video Object Removal
por: Zi, Bojia, et al.
Publicado: (2025)
por: Zi, Bojia, et al.
Publicado: (2025)
Fitted Q-Iteration via Max-Plus-Linear Approximation
por: Liu, Y., et al.
Publicado: (2024)
por: Liu, Y., et al.
Publicado: (2024)
FM-Agent: Scaling Formal Methods to Large Systems via LLM-Based Hoare-Style Reasoning
por: Ding, Haoran, et al.
Publicado: (2026)
por: Ding, Haoran, et al.
Publicado: (2026)
Ejemplares similares
-
The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence
por: MiniMax, et al.
Publicado: (2026) -
Discretizing Continuous Action Space with Unimodal Probability Distributions for On-Policy Reinforcement Learning
por: Zhu, Yuanyang, et al.
Publicado: (2024) -
A Hierarchical Deep Reinforcement Learning Framework for 6-DOF UCAV Air-to-Air Combat
por: Chai, Jiajun, et al.
Publicado: (2022) -
Computing Ex Ante Equilibrium in Heterogeneous Zero-Sum Team Games
por: Liu, Naming, et al.
Publicado: (2024) -
Bilevel Optimization over Saddle Points of Zero-Sum Markov Games
por: Zheng, Zihao, et al.
Publicado: (2026)