SPACeR: Self-Play Anchoring with Centralized Reference Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chang, Wei-Jer, Rangesh, Akshay, Joseph, Kevin, Strong, Matthew, Tomizuka, Masayoshi, Hu, Yihan, Zhan, Wei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914348997279744
author Chang, Wei-Jer
Rangesh, Akshay
Joseph, Kevin
Strong, Matthew
Tomizuka, Masayoshi
Hu, Yihan
Zhan, Wei
author_facet Chang, Wei-Jer
Rangesh, Akshay
Joseph, Kevin
Strong, Matthew
Tomizuka, Masayoshi
Hu, Yihan
Zhan, Wei
contents Developing autonomous vehicles (AVs) requires not only safety and efficiency, but also realistic, human-like behaviors that are socially aware and predictable. Achieving this requires sim agent policies that are human-like, fast, and scalable in multi-agent settings. Recent progress in imitation learning with large diffusion-based or tokenized models has shown that behaviors can be captured directly from human driving data, producing realistic policies. However, these models are computationally expensive, slow during inference, and struggle to adapt in reactive, closed-loop scenarios. In contrast, self-play reinforcement learning (RL) scales efficiently and naturally captures multi-agent interactions, but it often relies on heuristics and reward shaping, and the resulting policies can diverge from human norms. We propose SPACeR, a framework that leverages a pretrained tokenized autoregressive motion model as a centralized reference policy to guide decentralized self-play. The reference model provides likelihood rewards and KL divergence, anchoring policies to the human driving distribution while preserving RL scalability. Evaluated on the Waymo Sim Agents Challenge, our method achieves competitive performance with imitation-learned policies while being up to 10x faster at inference and 50x smaller in parameter size than large generative models. In addition, we demonstrate in closed-loop ego planning evaluation tasks that our sim agents can effectively measure planner quality with fast and scalable traffic simulation, establishing a new paradigm for testing autonomous driving policies.
format Preprint
id arxiv_https___arxiv_org_abs_2510_18060
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SPACeR: Self-Play Anchoring with Centralized Reference Models
Chang, Wei-Jer
Rangesh, Akshay
Joseph, Kevin
Strong, Matthew
Tomizuka, Masayoshi
Hu, Yihan
Zhan, Wei
Machine Learning
Artificial Intelligence
Robotics
I.2.9; I.2.6
Developing autonomous vehicles (AVs) requires not only safety and efficiency, but also realistic, human-like behaviors that are socially aware and predictable. Achieving this requires sim agent policies that are human-like, fast, and scalable in multi-agent settings. Recent progress in imitation learning with large diffusion-based or tokenized models has shown that behaviors can be captured directly from human driving data, producing realistic policies. However, these models are computationally expensive, slow during inference, and struggle to adapt in reactive, closed-loop scenarios. In contrast, self-play reinforcement learning (RL) scales efficiently and naturally captures multi-agent interactions, but it often relies on heuristics and reward shaping, and the resulting policies can diverge from human norms. We propose SPACeR, a framework that leverages a pretrained tokenized autoregressive motion model as a centralized reference policy to guide decentralized self-play. The reference model provides likelihood rewards and KL divergence, anchoring policies to the human driving distribution while preserving RL scalability. Evaluated on the Waymo Sim Agents Challenge, our method achieves competitive performance with imitation-learned policies while being up to 10x faster at inference and 50x smaller in parameter size than large generative models. In addition, we demonstrate in closed-loop ego planning evaluation tasks that our sim agents can effectively measure planner quality with fast and scalable traffic simulation, establishing a new paradigm for testing autonomous driving policies.
title SPACeR: Self-Play Anchoring with Centralized Reference Models
topic Machine Learning
Artificial Intelligence
Robotics
I.2.9; I.2.6
url https://arxiv.org/abs/2510.18060