PipeFlow: Pipelined Processing and Motion-Aware Frame Selection for Long-Form Video Editing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Munir, Mustafa, Rahman, Md Mostafijur, Bhardwaj, Kartikeya, Whatmough, Paul, Marculescu, Radu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914226922061824
author Munir, Mustafa
Rahman, Md Mostafijur
Bhardwaj, Kartikeya
Whatmough, Paul
Marculescu, Radu
author_facet Munir, Mustafa
Rahman, Md Mostafijur
Bhardwaj, Kartikeya
Whatmough, Paul
Marculescu, Radu
contents Long-form video editing poses unique challenges due to the exponential increase in the computational cost from joint editing and Denoising Diffusion Implicit Models (DDIM) inversion across extended sequences. To address these limitations, we propose PipeFlow, a scalable, pipelined video editing method that introduces three key innovations: First, based on a motion analysis using Structural Similarity Index Measure (SSIM) and Optical Flow, we identify and propose to skip editing of frames with low motion. Second, we propose a pipelined task scheduling algorithm that splits a video into multiple segments and performs DDIM inversion and joint editing in parallel based on available GPU memory. Lastly, we leverage a neural network-based interpolation technique to smooth out the border frames between segments and interpolate the previously skipped frames. Our method uniquely scales to longer videos by dividing them into smaller segments, allowing PipeFlow's editing time to increase linearly with video length. In principle, this enables editing of infinitely long videos without the growing per-frame computational overhead encountered by other methods. PipeFlow achieves up to a 9.6X speedup compared to TokenFlow and a 31.7X speedup over Diffusion Motion Transfer (DMT).
format Preprint
id arxiv_https___arxiv_org_abs_2512_24026
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PipeFlow: Pipelined Processing and Motion-Aware Frame Selection for Long-Form Video Editing
Munir, Mustafa
Rahman, Md Mostafijur
Bhardwaj, Kartikeya
Whatmough, Paul
Marculescu, Radu
Computer Vision and Pattern Recognition
Artificial Intelligence
Long-form video editing poses unique challenges due to the exponential increase in the computational cost from joint editing and Denoising Diffusion Implicit Models (DDIM) inversion across extended sequences. To address these limitations, we propose PipeFlow, a scalable, pipelined video editing method that introduces three key innovations: First, based on a motion analysis using Structural Similarity Index Measure (SSIM) and Optical Flow, we identify and propose to skip editing of frames with low motion. Second, we propose a pipelined task scheduling algorithm that splits a video into multiple segments and performs DDIM inversion and joint editing in parallel based on available GPU memory. Lastly, we leverage a neural network-based interpolation technique to smooth out the border frames between segments and interpolate the previously skipped frames. Our method uniquely scales to longer videos by dividing them into smaller segments, allowing PipeFlow's editing time to increase linearly with video length. In principle, this enables editing of infinitely long videos without the growing per-frame computational overhead encountered by other methods. PipeFlow achieves up to a 9.6X speedup compared to TokenFlow and a 31.7X speedup over Diffusion Motion Transfer (DMT).
title PipeFlow: Pipelined Processing and Motion-Aware Frame Selection for Long-Form Video Editing
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2512.24026