SoccerDiffusion: Toward Learning End-to-End Humanoid Robot Soccer from Gameplay Recordings

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Vahl, Florian, Griepenburg, Jörn, Gutsche, Jan, Güldenstein, Jasper, Zhang, Jianwei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908431015739392
author Vahl, Florian
Griepenburg, Jörn
Gutsche, Jan
Güldenstein, Jasper
Zhang, Jianwei
author_facet Vahl, Florian
Griepenburg, Jörn
Gutsche, Jan
Güldenstein, Jasper
Zhang, Jianwei
contents This paper introduces SoccerDiffusion, a transformer-based diffusion model designed to learn end-to-end control policies for humanoid robot soccer directly from real-world gameplay recordings. Using data collected from RoboCup competitions, the model predicts joint command trajectories from multi-modal sensor inputs, including vision, proprioception, and game state. We employ a distillation technique to enable real-time inference on embedded platforms that reduces the multi-step diffusion process to a single step. Our results demonstrate the model's ability to replicate complex motion behaviors such as walking, kicking, and fall recovery both in simulation and on physical robots. Although high-level tactical behavior remains limited, this work provides a robust foundation for subsequent reinforcement learning or preference optimization methods. We release the dataset, pretrained models, and code under: https://bit-bots.github.io/SoccerDiffusion
format Preprint
id arxiv_https___arxiv_org_abs_2504_20808
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SoccerDiffusion: Toward Learning End-to-End Humanoid Robot Soccer from Gameplay Recordings
Vahl, Florian
Griepenburg, Jörn
Gutsche, Jan
Güldenstein, Jasper
Zhang, Jianwei
Robotics
Artificial Intelligence
Machine Learning
This paper introduces SoccerDiffusion, a transformer-based diffusion model designed to learn end-to-end control policies for humanoid robot soccer directly from real-world gameplay recordings. Using data collected from RoboCup competitions, the model predicts joint command trajectories from multi-modal sensor inputs, including vision, proprioception, and game state. We employ a distillation technique to enable real-time inference on embedded platforms that reduces the multi-step diffusion process to a single step. Our results demonstrate the model's ability to replicate complex motion behaviors such as walking, kicking, and fall recovery both in simulation and on physical robots. Although high-level tactical behavior remains limited, this work provides a robust foundation for subsequent reinforcement learning or preference optimization methods. We release the dataset, pretrained models, and code under: https://bit-bots.github.io/SoccerDiffusion
title SoccerDiffusion: Toward Learning End-to-End Humanoid Robot Soccer from Gameplay Recordings
topic Robotics
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2504.20808