Masked Temporal Interpolation Diffusion for Procedure Planning in Instructional Videos

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Yufan, Qi, Zhaobo, Lin, Lingshuai, Jing, Junqi, Chai, Tingting, Zhang, Beichen, Wang, Shuhui, Zhang, Weigang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916825774686208
author Zhou, Yufan
Qi, Zhaobo
Lin, Lingshuai
Jing, Junqi
Chai, Tingting
Zhang, Beichen
Wang, Shuhui
Zhang, Weigang
author_facet Zhou, Yufan
Qi, Zhaobo
Lin, Lingshuai
Jing, Junqi
Chai, Tingting
Zhang, Beichen
Wang, Shuhui
Zhang, Weigang
contents In this paper, we address the challenge of procedure planning in instructional videos, aiming to generate coherent and task-aligned action sequences from start and end visual observations. Previous work has mainly relied on text-level supervision to bridge the gap between observed states and unobserved actions, but it struggles with capturing intricate temporal relationships among actions. Building on these efforts, we propose the Masked Temporal Interpolation Diffusion (MTID) model that introduces a latent space temporal interpolation module within the diffusion model. This module leverages a learnable interpolation matrix to generate intermediate latent features, thereby augmenting visual supervision with richer mid-state details. By integrating this enriched supervision into the model, we enable end-to-end training tailored to task-specific requirements, significantly enhancing the model's capacity to predict temporally coherent action sequences. Additionally, we introduce an action-aware mask projection mechanism to restrict the action generation space, combined with a task-adaptive masked proximity loss to prioritize more accurate reasoning results close to the given start and end states over those in intermediate steps. Simultaneously, it filters out task-irrelevant action predictions, leading to contextually aware action sequences. Experimental results across three widely used benchmark datasets demonstrate that our MTID achieves promising action planning performance on most metrics. The code is available at https://github.com/WiserZhou/MTID.
format Preprint
id arxiv_https___arxiv_org_abs_2507_03393
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Masked Temporal Interpolation Diffusion for Procedure Planning in Instructional Videos
Zhou, Yufan
Qi, Zhaobo
Lin, Lingshuai
Jing, Junqi
Chai, Tingting
Zhang, Beichen
Wang, Shuhui
Zhang, Weigang
Computer Vision and Pattern Recognition
In this paper, we address the challenge of procedure planning in instructional videos, aiming to generate coherent and task-aligned action sequences from start and end visual observations. Previous work has mainly relied on text-level supervision to bridge the gap between observed states and unobserved actions, but it struggles with capturing intricate temporal relationships among actions. Building on these efforts, we propose the Masked Temporal Interpolation Diffusion (MTID) model that introduces a latent space temporal interpolation module within the diffusion model. This module leverages a learnable interpolation matrix to generate intermediate latent features, thereby augmenting visual supervision with richer mid-state details. By integrating this enriched supervision into the model, we enable end-to-end training tailored to task-specific requirements, significantly enhancing the model's capacity to predict temporally coherent action sequences. Additionally, we introduce an action-aware mask projection mechanism to restrict the action generation space, combined with a task-adaptive masked proximity loss to prioritize more accurate reasoning results close to the given start and end states over those in intermediate steps. Simultaneously, it filters out task-irrelevant action predictions, leading to contextually aware action sequences. Experimental results across three widely used benchmark datasets demonstrate that our MTID achieves promising action planning performance on most metrics. The code is available at https://github.com/WiserZhou/MTID.
title Masked Temporal Interpolation Diffusion for Procedure Planning in Instructional Videos
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.03393