Frame Guidance: Training-Free Guidance for Frame-Level Control in Video Diffusion Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Jang, Sangwon, Ki, Taekyung, Jo, Jaehyeong, Yoon, Jaehong, Kim, Soo Ye, Lin, Zhe, Hwang, Sung Ju
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915829381070848
author Jang, Sangwon
Ki, Taekyung
Jo, Jaehyeong
Yoon, Jaehong
Kim, Soo Ye
Lin, Zhe
Hwang, Sung Ju
author_facet Jang, Sangwon
Ki, Taekyung
Jo, Jaehyeong
Yoon, Jaehong
Kim, Soo Ye
Lin, Zhe
Hwang, Sung Ju
contents Advancements in diffusion models have significantly improved video quality, directing attention to fine-grained controllability. However, many existing methods depend on fine-tuning large-scale video models for specific tasks, which becomes increasingly impractical as model sizes continue to grow. In this work, we present Frame Guidance, a training-free guidance for controllable video generation based on frame-level signals, such as keyframes, style reference images, sketches, or depth maps. For practical training-free guidance, we propose a simple latent processing method that dramatically reduces memory usage, and apply a novel latent optimization strategy designed for globally coherent video generation. Frame Guidance enables effective control across diverse tasks, including keyframe guidance, stylization, and looping, without any training, compatible with any video models. Experimental results show that Frame Guidance can produce high-quality controlled videos for a wide range of tasks and input signals.
format Preprint
id arxiv_https___arxiv_org_abs_2506_07177
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Frame Guidance: Training-Free Guidance for Frame-Level Control in Video Diffusion Models
Jang, Sangwon
Ki, Taekyung
Jo, Jaehyeong
Yoon, Jaehong
Kim, Soo Ye
Lin, Zhe
Hwang, Sung Ju
Computer Vision and Pattern Recognition
Artificial Intelligence
Advancements in diffusion models have significantly improved video quality, directing attention to fine-grained controllability. However, many existing methods depend on fine-tuning large-scale video models for specific tasks, which becomes increasingly impractical as model sizes continue to grow. In this work, we present Frame Guidance, a training-free guidance for controllable video generation based on frame-level signals, such as keyframes, style reference images, sketches, or depth maps. For practical training-free guidance, we propose a simple latent processing method that dramatically reduces memory usage, and apply a novel latent optimization strategy designed for globally coherent video generation. Frame Guidance enables effective control across diverse tasks, including keyframe guidance, stylization, and looping, without any training, compatible with any video models. Experimental results show that Frame Guidance can produce high-quality controlled videos for a wide range of tasks and input signals.
title Frame Guidance: Training-Free Guidance for Frame-Level Control in Video Diffusion Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2506.07177