Φ-Noise: Training-Free Temporal Video Conditioning via Phase-Based Noise Manipulation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Abramovich, Ofir, Cohen, Nadav Z., Rosenthal, Adi, Shamir, Ariel
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918519892869120
author Abramovich, Ofir
Cohen, Nadav Z.
Rosenthal, Adi
Shamir, Ariel
author_facet Abramovich, Ofir
Cohen, Nadav Z.
Rosenthal, Adi
Shamir, Ariel
contents Latent video diffusion models generate videos by progressively transforming Gaussian noise into realistic samples conditioned on text or visual inputs. However, existing conditioning methods often require additional training and computational overhead. Motivated by recent findings on the importance of frequency components in generative models, we propose a simple, training-free approach for motion-conditioned video generation by injecting low-frequency phase information from a reference video directly into the diffusion noise latents. Our method transfers motion cues without modifying the model architecture or inference pipeline. Using several applications, we demonstrate effective control over both appearance and dynamics in generated videos, while achieving competitive or superior results compared to more complex conditioning approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2605_24509
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Φ-Noise: Training-Free Temporal Video Conditioning via Phase-Based Noise Manipulation
Abramovich, Ofir
Cohen, Nadav Z.
Rosenthal, Adi
Shamir, Ariel
Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
Machine Learning
Latent video diffusion models generate videos by progressively transforming Gaussian noise into realistic samples conditioned on text or visual inputs. However, existing conditioning methods often require additional training and computational overhead. Motivated by recent findings on the importance of frequency components in generative models, we propose a simple, training-free approach for motion-conditioned video generation by injecting low-frequency phase information from a reference video directly into the diffusion noise latents. Our method transfers motion cues without modifying the model architecture or inference pipeline. Using several applications, we demonstrate effective control over both appearance and dynamics in generated videos, while achieving competitive or superior results compared to more complex conditioning approaches.
title Φ-Noise: Training-Free Temporal Video Conditioning via Phase-Based Noise Manipulation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
Machine Learning
url https://arxiv.org/abs/2605.24509