Superhuman AI for Stratego Using Self-Play Reinforcement Learning and Test-Time Search
Fuente:
arXiv
Salvato in:
| Autori principali: | Sokota, Samuel, Vinitsky, Eugene, Hu, Hengyuan, Kolter, J. Zico, Farina, Gabriele |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The Update-Equivalence Framework for Decision-Time Planning
di: Sokota, Samuel, et al.
Pubblicazione: (2023)
di: Sokota, Samuel, et al.
Pubblicazione: (2023)
Reevaluating Policy Gradient Methods for Imperfect-Information Games
di: Rudolph, Max, et al.
Pubblicazione: (2025)
di: Rudolph, Max, et al.
Pubblicazione: (2025)
Test-Time Adaptation Induces Stronger Accuracy and Agreement-on-the-Line
di: Kim, Eungyeup, et al.
Pubblicazione: (2023)
di: Kim, Eungyeup, et al.
Pubblicazione: (2023)
Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning
di: Xu, Yixuan Even, et al.
Pubblicazione: (2025)
di: Xu, Yixuan Even, et al.
Pubblicazione: (2025)
Scalable Oversight for Superhuman AI via Recursive Self-Critiquing
di: Wen, Xueru, et al.
Pubblicazione: (2025)
di: Wen, Xueru, et al.
Pubblicazione: (2025)
Human-compatible driving partners through data-regularized self-play reinforcement learning
di: Cornelisse, Daphne, et al.
Pubblicazione: (2024)
di: Cornelisse, Daphne, et al.
Pubblicazione: (2024)
Mimetic Initialization of MLPs
di: Trockman, Asher, et al.
Pubblicazione: (2026)
di: Trockman, Asher, et al.
Pubblicazione: (2026)
Imitation Bootstrapped Reinforcement Learning
di: Hu, Hengyuan, et al.
Pubblicazione: (2023)
di: Hu, Hengyuan, et al.
Pubblicazione: (2023)
Self-Harmony: Learning to Harmonize Self-Supervision and Self-Play in Test-Time Reinforcement Learning
di: Wang, Ru, et al.
Pubblicazione: (2025)
di: Wang, Ru, et al.
Pubblicazione: (2025)
Video Game Level Design as a Multi-Agent Reinforcement Learning Problem
di: Earle, Sam, et al.
Pubblicazione: (2025)
di: Earle, Sam, et al.
Pubblicazione: (2025)
Robust Autonomy Emerges from Self-Play
di: Cusumano-Towner, Marco, et al.
Pubblicazione: (2025)
di: Cusumano-Towner, Marco, et al.
Pubblicazione: (2025)
Superhuman Safe and Agile Racing through Multi-Agent Reinforcement Learning
di: Geles, Ismail, et al.
Pubblicazione: (2026)
di: Geles, Ismail, et al.
Pubblicazione: (2026)
Weight Ensembling Improves Reasoning in Language Models
di: Dang, Xingyu, et al.
Pubblicazione: (2025)
di: Dang, Xingyu, et al.
Pubblicazione: (2025)
Base Models Look Human To AI Detectors
di: Xu, Yixuan Even, et al.
Pubblicazione: (2026)
di: Xu, Yixuan Even, et al.
Pubblicazione: (2026)
Open-Endedness is Essential for Artificial Superhuman Intelligence
di: Hughes, Edward, et al.
Pubblicazione: (2024)
di: Hughes, Edward, et al.
Pubblicazione: (2024)
Extending Test-Time Scaling: A 3D Perspective with Context, Batch, and Turn
di: Yu, Chao, et al.
Pubblicazione: (2025)
di: Yu, Chao, et al.
Pubblicazione: (2025)
A Simple and Effective Pruning Approach for Large Language Models
di: Sun, Mingjie, et al.
Pubblicazione: (2023)
di: Sun, Mingjie, et al.
Pubblicazione: (2023)
Looking beyond the next token
di: Thankaraj, Abitha, et al.
Pubblicazione: (2025)
di: Thankaraj, Abitha, et al.
Pubblicazione: (2025)
CoSPlay: Cooperative Self-Play at Test-Time with Self-Generated Code and Unit Test
di: Hu, Zhangyi, et al.
Pubblicazione: (2026)
di: Hu, Zhangyi, et al.
Pubblicazione: (2026)
Provably Bounding Neural Network Preimages
di: Kotha, Suhas, et al.
Pubblicazione: (2023)
di: Kotha, Suhas, et al.
Pubblicazione: (2023)
Understanding Optimization in Deep Learning with Central Flows
di: Cohen, Jeremy M., et al.
Pubblicazione: (2024)
di: Cohen, Jeremy M., et al.
Pubblicazione: (2024)
Transformers to SSMs: Distilling Quadratic Knowledge to Subquadratic Models
di: Bick, Aviv, et al.
Pubblicazione: (2024)
di: Bick, Aviv, et al.
Pubblicazione: (2024)
Accelerating Diffusion Planners in Offline RL via Reward-Aware Consistency Trajectory Distillation
di: Duan, Xintong, et al.
Pubblicazione: (2025)
di: Duan, Xintong, et al.
Pubblicazione: (2025)
Neural Network Verification with Branch-and-Bound for General Nonlinearities
di: Shi, Zhouxing, et al.
Pubblicazione: (2024)
di: Shi, Zhouxing, et al.
Pubblicazione: (2024)
Contextures: Representations from Contexts
di: Zhai, Runtian, et al.
Pubblicazione: (2025)
di: Zhai, Runtian, et al.
Pubblicazione: (2025)
TriPlay-RL: Tri-Role Self-Play Reinforcement Learning for LLM Safety Alignment
di: Tan, Zhewen, et al.
Pubblicazione: (2026)
di: Tan, Zhewen, et al.
Pubblicazione: (2026)
ROGUE: Misaligned Agent Behavior Arising from Ordinary Computer Use
di: Tien, Jeremy, et al.
Pubblicazione: (2026)
di: Tien, Jeremy, et al.
Pubblicazione: (2026)
Self-Play Reinforcement Learning under Imperfect Information in Big 2
di: Patwa, Aalok
Pubblicazione: (2026)
di: Patwa, Aalok
Pubblicazione: (2026)
ICPL: Few-shot In-context Preference Learning via LLMs
di: Yu, Chao, et al.
Pubblicazione: (2024)
di: Yu, Chao, et al.
Pubblicazione: (2024)
Self-Improving AI Agents through Self-Play
di: Chojecki, Przemyslaw
Pubblicazione: (2025)
di: Chojecki, Przemyslaw
Pubblicazione: (2025)
Lyapunov-Guided Self-Alignment: Test-Time Adaptation for Offline Safe Reinforcement Learning
di: Han, Seungyub, et al.
Pubblicazione: (2026)
di: Han, Seungyub, et al.
Pubblicazione: (2026)
Plug-and-Play Transformer Modules for Test-Time Adaptation
di: Chang, Xiangyu, et al.
Pubblicazione: (2024)
di: Chang, Xiangyu, et al.
Pubblicazione: (2024)
Inference Optimal VLMs Need Fewer Visual Tokens and More Parameters
di: Li, Kevin Y., et al.
Pubblicazione: (2024)
di: Li, Kevin Y., et al.
Pubblicazione: (2024)
Compute-Optimal LLMs Provably Generalize Better With Scale
di: Finzi, Marc, et al.
Pubblicazione: (2025)
di: Finzi, Marc, et al.
Pubblicazione: (2025)
Differentially Private Reinforcement Learning with Self-Play
di: Qiao, Dan, et al.
Pubblicazione: (2024)
di: Qiao, Dan, et al.
Pubblicazione: (2024)
Learning on the Job: Test-Time Curricula for Targeted Reinforcement Learning
di: Hübotter, Jonas, et al.
Pubblicazione: (2025)
di: Hübotter, Jonas, et al.
Pubblicazione: (2025)
GAE Falls Short in Imperfect-Information Self-Play Reinforcement Learning
di: Fan, Zhiyuan, et al.
Pubblicazione: (2026)
di: Fan, Zhiyuan, et al.
Pubblicazione: (2026)
Near-Optimal Reinforcement Learning with Self-Play under Adaptivity Constraints
di: Qiao, Dan, et al.
Pubblicazione: (2024)
di: Qiao, Dan, et al.
Pubblicazione: (2024)
AceSearcher: Bootstrapping Reasoning and Search for LLMs via Reinforced Self-Play
di: Xu, Ran, et al.
Pubblicazione: (2025)
di: Xu, Ran, et al.
Pubblicazione: (2025)
One-Step Diffusion Distillation through Score Implicit Matching
di: Luo, Weijian, et al.
Pubblicazione: (2024)
di: Luo, Weijian, et al.
Pubblicazione: (2024)
Documenti analoghi
-
The Update-Equivalence Framework for Decision-Time Planning
di: Sokota, Samuel, et al.
Pubblicazione: (2023) -
Reevaluating Policy Gradient Methods for Imperfect-Information Games
di: Rudolph, Max, et al.
Pubblicazione: (2025) -
Test-Time Adaptation Induces Stronger Accuracy and Agreement-on-the-Line
di: Kim, Eungyeup, et al.
Pubblicazione: (2023) -
Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning
di: Xu, Yixuan Even, et al.
Pubblicazione: (2025) -
Scalable Oversight for Superhuman AI via Recursive Self-Critiquing
di: Wen, Xueru, et al.
Pubblicazione: (2025)