Closing the Train-Test Gap in World Models for Gradient-Based Planning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Parthasarathy, Arjun, Kalra, Nimit, Agrawal, Rohun, LeCun, Yann, Bounou, Oumayma, Izmailov, Pavel, Goldblum, Micah
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914193037328384
author Parthasarathy, Arjun
Kalra, Nimit
Agrawal, Rohun
LeCun, Yann
Bounou, Oumayma
Izmailov, Pavel
Goldblum, Micah
author_facet Parthasarathy, Arjun
Kalra, Nimit
Agrawal, Rohun
LeCun, Yann
Bounou, Oumayma
Izmailov, Pavel
Goldblum, Micah
contents World models paired with model predictive control (MPC) can be trained offline on large-scale datasets of expert trajectories and enable generalization to a wide range of planning tasks at inference time. Compared to traditional MPC procedures, which rely on slow search algorithms or on iteratively solving optimization problems exactly, gradient-based planning offers a computationally efficient alternative. However, the performance of gradient-based planning has thus far lagged behind that of other approaches. In this paper, we propose improved methods for training world models that enable efficient gradient-based planning. We begin with the observation that although a world model is trained on a next-state prediction objective, it is used at test-time to instead estimate a sequence of actions. The goal of our work is to close this train-test gap. To that end, we propose train-time data synthesis techniques that enable significantly improved gradient-based planning with existing world models. At test time, our approach outperforms or matches the classical gradient-free cross-entropy method (CEM) across a variety of object manipulation and navigation tasks in 10% of the time budget.
format Preprint
id arxiv_https___arxiv_org_abs_2512_09929
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Closing the Train-Test Gap in World Models for Gradient-Based Planning
Parthasarathy, Arjun
Kalra, Nimit
Agrawal, Rohun
LeCun, Yann
Bounou, Oumayma
Izmailov, Pavel
Goldblum, Micah
Machine Learning
Robotics
World models paired with model predictive control (MPC) can be trained offline on large-scale datasets of expert trajectories and enable generalization to a wide range of planning tasks at inference time. Compared to traditional MPC procedures, which rely on slow search algorithms or on iteratively solving optimization problems exactly, gradient-based planning offers a computationally efficient alternative. However, the performance of gradient-based planning has thus far lagged behind that of other approaches. In this paper, we propose improved methods for training world models that enable efficient gradient-based planning. We begin with the observation that although a world model is trained on a next-state prediction objective, it is used at test-time to instead estimate a sequence of actions. The goal of our work is to close this train-test gap. To that end, we propose train-time data synthesis techniques that enable significantly improved gradient-based planning with existing world models. At test time, our approach outperforms or matches the classical gradient-free cross-entropy method (CEM) across a variety of object manipulation and navigation tasks in 10% of the time budget.
title Closing the Train-Test Gap in World Models for Gradient-Based Planning
topic Machine Learning
Robotics
url https://arxiv.org/abs/2512.09929