Adapting a World Model for Trajectory Following in a 3D Game

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tot, Marko, Ishida, Shu, Lemkhenter, Abdelhak, Bignell, David, Choudhury, Pallavi, Lovett, Chris, França, Luis, de Mendonça, Matheus Ribeiro Furtado, Gupta, Tarun, Gehring, Darren, Devlin, Sam, Macua, Sergio Valcarcel, Georgescu, Raluca
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913796622123008
author Tot, Marko
Ishida, Shu
Lemkhenter, Abdelhak
Bignell, David
Choudhury, Pallavi
Lovett, Chris
França, Luis
de Mendonça, Matheus Ribeiro Furtado
Gupta, Tarun
Gehring, Darren
Devlin, Sam
Macua, Sergio Valcarcel
Georgescu, Raluca
author_facet Tot, Marko
Ishida, Shu
Lemkhenter, Abdelhak
Bignell, David
Choudhury, Pallavi
Lovett, Chris
França, Luis
de Mendonça, Matheus Ribeiro Furtado
Gupta, Tarun
Gehring, Darren
Devlin, Sam
Macua, Sergio Valcarcel
Georgescu, Raluca
contents Imitation learning is a powerful tool for training agents by leveraging expert knowledge, and being able to replicate a given trajectory is an integral part of it. In complex environments, like modern 3D video games, distribution shift and stochasticity necessitate robust approaches beyond simple action replay. In this study, we apply Inverse Dynamics Models (IDM) with different encoders and policy heads to trajectory following in a modern 3D video game -- Bleeding Edge. Additionally, we investigate several future alignment strategies that address the distribution shift caused by the aleatoric uncertainty and imperfections of the agent. We measure both the trajectory deviation distance and the first significant deviation point between the reference and the agent's trajectory and show that the optimal configuration depends on the chosen setting. Our results show that in a diverse data setting, a GPT-style policy head with an encoder trained from scratch performs the best, DINOv2 encoder with the GPT-style policy head gives the best results in the low data regime, and both GPT-style and MLP-style policy heads had comparable results when pre-trained on a diverse setting and fine-tuned for a specific behaviour setting.
format Preprint
id arxiv_https___arxiv_org_abs_2504_12299
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Adapting a World Model for Trajectory Following in a 3D Game
Tot, Marko
Ishida, Shu
Lemkhenter, Abdelhak
Bignell, David
Choudhury, Pallavi
Lovett, Chris
França, Luis
de Mendonça, Matheus Ribeiro Furtado
Gupta, Tarun
Gehring, Darren
Devlin, Sam
Macua, Sergio Valcarcel
Georgescu, Raluca
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
Imitation learning is a powerful tool for training agents by leveraging expert knowledge, and being able to replicate a given trajectory is an integral part of it. In complex environments, like modern 3D video games, distribution shift and stochasticity necessitate robust approaches beyond simple action replay. In this study, we apply Inverse Dynamics Models (IDM) with different encoders and policy heads to trajectory following in a modern 3D video game -- Bleeding Edge. Additionally, we investigate several future alignment strategies that address the distribution shift caused by the aleatoric uncertainty and imperfections of the agent. We measure both the trajectory deviation distance and the first significant deviation point between the reference and the agent's trajectory and show that the optimal configuration depends on the chosen setting. Our results show that in a diverse data setting, a GPT-style policy head with an encoder trained from scratch performs the best, DINOv2 encoder with the GPT-style policy head gives the best results in the low data regime, and both GPT-style and MLP-style policy heads had comparable results when pre-trained on a diverse setting and fine-tuned for a specific behaviour setting.
title Adapting a World Model for Trajectory Following in a 3D Game
topic Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2504.12299