Image Diffusion Models Exhibit Emergent Temporal Propagation in Videos

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Kim, Youngseo, Kim, Dohyun, Han, Geonhee, Seo, Paul Hongsuck
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909924112465920
author Kim, Youngseo
Kim, Dohyun
Han, Geonhee
Seo, Paul Hongsuck
author_facet Kim, Youngseo
Kim, Dohyun
Han, Geonhee
Seo, Paul Hongsuck
contents Image diffusion models, though originally developed for image generation, implicitly capture rich semantic structures that enable various recognition and localization tasks beyond synthesis. In this work, we investigate their self-attention maps can be reinterpreted as semantic label propagation kernels, providing robust pixel-level correspondences between relevant image regions. Extending this mechanism across frames yields a temporal propagation kernel that enables zero-shot object tracking via segmentation in videos. We further demonstrate the effectiveness of test-time optimization strategies-DDIM inversion, textual inversion, and adaptive head weighting-in adapting diffusion features for robust and consistent label propagation. Building on these findings, we introduce DRIFT, a framework for object tracking in videos leveraging a pretrained image diffusion model with SAM-guided mask refinement, achieving state-of-the-art zero-shot performance on standard video object segmentation benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2511_19936
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Image Diffusion Models Exhibit Emergent Temporal Propagation in Videos
Kim, Youngseo
Kim, Dohyun
Han, Geonhee
Seo, Paul Hongsuck
Computer Vision and Pattern Recognition
Image diffusion models, though originally developed for image generation, implicitly capture rich semantic structures that enable various recognition and localization tasks beyond synthesis. In this work, we investigate their self-attention maps can be reinterpreted as semantic label propagation kernels, providing robust pixel-level correspondences between relevant image regions. Extending this mechanism across frames yields a temporal propagation kernel that enables zero-shot object tracking via segmentation in videos. We further demonstrate the effectiveness of test-time optimization strategies-DDIM inversion, textual inversion, and adaptive head weighting-in adapting diffusion features for robust and consistent label propagation. Building on these findings, we introduce DRIFT, a framework for object tracking in videos leveraging a pretrained image diffusion model with SAM-guided mask refinement, achieving state-of-the-art zero-shot performance on standard video object segmentation benchmarks.
title Image Diffusion Models Exhibit Emergent Temporal Propagation in Videos
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.19936