Φ-Noise: Training-Free Temporal Video Conditioning via Phase-Based Noise Manipulation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866918519892869120 |
|---|---|
| author | Abramovich, Ofir Cohen, Nadav Z. Rosenthal, Adi Shamir, Ariel |
| author_facet | Abramovich, Ofir Cohen, Nadav Z. Rosenthal, Adi Shamir, Ariel |
| contents | Latent video diffusion models generate videos by progressively transforming Gaussian noise into realistic samples conditioned on text or visual inputs. However, existing conditioning methods often require additional training and computational overhead. Motivated by recent findings on the importance of frequency components in generative models, we propose a simple, training-free approach for motion-conditioned video generation by injecting low-frequency phase information from a reference video directly into the diffusion noise latents. Our method transfers motion cues without modifying the model architecture or inference pipeline. Using several applications, we demonstrate effective control over both appearance and dynamics in generated videos, while achieving competitive or superior results compared to more complex conditioning approaches. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_24509 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Φ-Noise: Training-Free Temporal Video Conditioning via Phase-Based Noise Manipulation Abramovich, Ofir Cohen, Nadav Z. Rosenthal, Adi Shamir, Ariel Computer Vision and Pattern Recognition Artificial Intelligence Graphics Machine Learning Latent video diffusion models generate videos by progressively transforming Gaussian noise into realistic samples conditioned on text or visual inputs. However, existing conditioning methods often require additional training and computational overhead. Motivated by recent findings on the importance of frequency components in generative models, we propose a simple, training-free approach for motion-conditioned video generation by injecting low-frequency phase information from a reference video directly into the diffusion noise latents. Our method transfers motion cues without modifying the model architecture or inference pipeline. Using several applications, we demonstrate effective control over both appearance and dynamics in generated videos, while achieving competitive or superior results compared to more complex conditioning approaches. |
| title | Φ-Noise: Training-Free Temporal Video Conditioning via Phase-Based Noise Manipulation |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Graphics Machine Learning |
| url | https://arxiv.org/abs/2605.24509 |