FreeAction: Training-Free Techniques for Enhanced Fidelity of Trajectory-to-Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Seungwook, Lee, Seunghyeon, Cho, Minsu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909814449242112
author Kim, Seungwook
Lee, Seunghyeon
Cho, Minsu
author_facet Kim, Seungwook
Lee, Seunghyeon
Cho, Minsu
contents Generating realistic robot videos from explicit action trajectories is a critical step toward building effective world models and robotics foundation models. We introduce two training-free, inference-time techniques that fully exploit explicit action parameters in diffusion-based robot video generation. Instead of treating action vectors as passive conditioning signals, our methods actively incorporate them to guide both the classifier-free guidance process and the initialization of Gaussian latents. First, action-scaled classifier-free guidance dynamically modulates guidance strength in proportion to action magnitude, enhancing controllability over motion intensity. Second, action-scaled noise truncation adjusts the distribution of initially sampled noise to better align with the desired motion dynamics. Experiments on real robot manipulation datasets demonstrate that these techniques significantly improve action coherence and visual quality across diverse robot environments.
format Preprint
id arxiv_https___arxiv_org_abs_2509_24241
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FreeAction: Training-Free Techniques for Enhanced Fidelity of Trajectory-to-Video Generation
Kim, Seungwook
Lee, Seunghyeon
Cho, Minsu
Computer Vision and Pattern Recognition
Robotics
Generating realistic robot videos from explicit action trajectories is a critical step toward building effective world models and robotics foundation models. We introduce two training-free, inference-time techniques that fully exploit explicit action parameters in diffusion-based robot video generation. Instead of treating action vectors as passive conditioning signals, our methods actively incorporate them to guide both the classifier-free guidance process and the initialization of Gaussian latents. First, action-scaled classifier-free guidance dynamically modulates guidance strength in proportion to action magnitude, enhancing controllability over motion intensity. Second, action-scaled noise truncation adjusts the distribution of initially sampled noise to better align with the desired motion dynamics. Experiments on real robot manipulation datasets demonstrate that these techniques significantly improve action coherence and visual quality across diverse robot environments.
title FreeAction: Training-Free Techniques for Enhanced Fidelity of Trajectory-to-Video Generation
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2509.24241