ActionParty: Multi-Subject Action Binding in Generative Video Games

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pondaven, Alexander, Wu, Ziyi, Gilitschenski, Igor, Torr, Philip, Tulyakov, Sergey, Pizzati, Fabio, Siarohin, Aliaksandr
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912999487307776
author Pondaven, Alexander
Wu, Ziyi
Gilitschenski, Igor
Torr, Philip
Tulyakov, Sergey
Pizzati, Fabio
Siarohin, Aliaksandr
author_facet Pondaven, Alexander
Wu, Ziyi
Gilitschenski, Igor
Torr, Philip
Tulyakov, Sergey
Pizzati, Fabio
Siarohin, Aliaksandr
contents Recent advances in video diffusion have enabled the development of "world models" capable of simulating interactive environments. However, these models are largely restricted to single-agent settings, failing to control multiple agents simultaneously in a scene. In this work, we tackle a fundamental issue of action binding in existing video diffusion models, which struggle to associate specific actions with their corresponding subjects. For this purpose, we propose ActionParty, an action controllable multi-subject world model for generative video games. It introduces subject state tokens, i.e. latent variables that persistently capture the state of each subject in the scene. By jointly modeling state tokens and video latents with a spatial biasing mechanism, we disentangle global video frame rendering from individual action-controlled subject updates. We evaluate ActionParty on the Melting Pot benchmark, demonstrating the first video world model capable of controlling up to seven players simultaneously across 46 diverse environments. Our results show significant improvements in action-following accuracy and identity consistency, while enabling robust autoregressive tracking of subjects through complex interactions.
format Preprint
id arxiv_https___arxiv_org_abs_2604_02330
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ActionParty: Multi-Subject Action Binding in Generative Video Games
Pondaven, Alexander
Wu, Ziyi
Gilitschenski, Igor
Torr, Philip
Tulyakov, Sergey
Pizzati, Fabio
Siarohin, Aliaksandr
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Recent advances in video diffusion have enabled the development of "world models" capable of simulating interactive environments. However, these models are largely restricted to single-agent settings, failing to control multiple agents simultaneously in a scene. In this work, we tackle a fundamental issue of action binding in existing video diffusion models, which struggle to associate specific actions with their corresponding subjects. For this purpose, we propose ActionParty, an action controllable multi-subject world model for generative video games. It introduces subject state tokens, i.e. latent variables that persistently capture the state of each subject in the scene. By jointly modeling state tokens and video latents with a spatial biasing mechanism, we disentangle global video frame rendering from individual action-controlled subject updates. We evaluate ActionParty on the Melting Pot benchmark, demonstrating the first video world model capable of controlling up to seven players simultaneously across 46 diverse environments. Our results show significant improvements in action-following accuracy and identity consistency, while enabling robust autoregressive tracking of subjects through complex interactions.
title ActionParty: Multi-Subject Action Binding in Generative Video Games
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2604.02330