TransFlow: Motion Knowledge Transfer from Video Diffusion Models to Video Salient Object Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cho, Suhwan, Lee, Minhyeok, Lee, Jungho, Yang, Sunghun, Lee, Sangyoun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911078348226560
author Cho, Suhwan
Lee, Minhyeok
Lee, Jungho
Yang, Sunghun
Lee, Sangyoun
author_facet Cho, Suhwan
Lee, Minhyeok
Lee, Jungho
Yang, Sunghun
Lee, Sangyoun
contents Video salient object detection (SOD) relies on motion cues to distinguish salient objects from backgrounds, but training such models is limited by scarce video datasets compared to abundant image datasets. Existing approaches that use spatial transformations to create video sequences from static images fail for motion-guided tasks, as these transformations produce unrealistic optical flows that lack semantic understanding of motion. We present TransFlow, which transfers motion knowledge from pre-trained video diffusion models to generate realistic training data for video SOD. Video diffusion models have learned rich semantic motion priors from large-scale video data, understanding how different objects naturally move in real scenes. TransFlow leverages this knowledge to generate semantically-aware optical flows from static images, where objects exhibit natural motion patterns while preserving spatial boundaries and temporal coherence. Our method achieves improved performance across multiple benchmarks, demonstrating effective motion knowledge transfer.
format Preprint
id arxiv_https___arxiv_org_abs_2507_19789
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TransFlow: Motion Knowledge Transfer from Video Diffusion Models to Video Salient Object Detection
Cho, Suhwan
Lee, Minhyeok
Lee, Jungho
Yang, Sunghun
Lee, Sangyoun
Computer Vision and Pattern Recognition
Video salient object detection (SOD) relies on motion cues to distinguish salient objects from backgrounds, but training such models is limited by scarce video datasets compared to abundant image datasets. Existing approaches that use spatial transformations to create video sequences from static images fail for motion-guided tasks, as these transformations produce unrealistic optical flows that lack semantic understanding of motion. We present TransFlow, which transfers motion knowledge from pre-trained video diffusion models to generate realistic training data for video SOD. Video diffusion models have learned rich semantic motion priors from large-scale video data, understanding how different objects naturally move in real scenes. TransFlow leverages this knowledge to generate semantically-aware optical flows from static images, where objects exhibit natural motion patterns while preserving spatial boundaries and temporal coherence. Our method achieves improved performance across multiple benchmarks, demonstrating effective motion knowledge transfer.
title TransFlow: Motion Knowledge Transfer from Video Diffusion Models to Video Salient Object Detection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.19789