SneakPeek: Future-Guided Instructional Streaming Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hong, Cheeun, Barquero, German, Sener, Fadime, Georgopoulos, Markos, Schönfeld, Edgar, Popov, Stefan, Du, Yuming, Mañas, Oscar, Pumarola, Albert
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917146759528448
author Hong, Cheeun
Barquero, German
Sener, Fadime
Georgopoulos, Markos
Schönfeld, Edgar
Popov, Stefan
Du, Yuming
Mañas, Oscar
Pumarola, Albert
author_facet Hong, Cheeun
Barquero, German
Sener, Fadime
Georgopoulos, Markos
Schönfeld, Edgar
Popov, Stefan
Du, Yuming
Mañas, Oscar
Pumarola, Albert
contents Instructional video generation is an emerging task that aims to synthesize coherent demonstrations of procedural activities from textual descriptions. Such capability has broad implications for content creation, education, and human-AI interaction, yet existing video diffusion models struggle to maintain temporal consistency and controllability across long sequences of multiple action steps. We introduce a pipeline for future-driven streaming instructional video generation, dubbed SneakPeek, a diffusion-based autoregressive framework designed to generate precise, stepwise instructional videos conditioned on an initial image and structured textual prompts. Our approach introduces three key innovations to enhance consistency and controllability: (1) predictive causal adaptation, where a causal model learns to perform next-frame prediction and anticipate future keyframes; (2) future-guided self-forcing with a dual-region KV caching scheme to address the exposure bias issue at inference time; (3) multi-prompt conditioning, which provides fine-grained and procedural control over multi-step instructions. Together, these components mitigate temporal drift, preserve motion consistency, and enable interactive video generation where future prompt updates dynamically influence ongoing streaming video generation. Experimental results demonstrate that our method produces temporally coherent and semantically faithful instructional videos that accurately follow complex, multi-step task descriptions.
format Preprint
id arxiv_https___arxiv_org_abs_2512_13019
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SneakPeek: Future-Guided Instructional Streaming Video Generation
Hong, Cheeun
Barquero, German
Sener, Fadime
Georgopoulos, Markos
Schönfeld, Edgar
Popov, Stefan
Du, Yuming
Mañas, Oscar
Pumarola, Albert
Computer Vision and Pattern Recognition
Instructional video generation is an emerging task that aims to synthesize coherent demonstrations of procedural activities from textual descriptions. Such capability has broad implications for content creation, education, and human-AI interaction, yet existing video diffusion models struggle to maintain temporal consistency and controllability across long sequences of multiple action steps. We introduce a pipeline for future-driven streaming instructional video generation, dubbed SneakPeek, a diffusion-based autoregressive framework designed to generate precise, stepwise instructional videos conditioned on an initial image and structured textual prompts. Our approach introduces three key innovations to enhance consistency and controllability: (1) predictive causal adaptation, where a causal model learns to perform next-frame prediction and anticipate future keyframes; (2) future-guided self-forcing with a dual-region KV caching scheme to address the exposure bias issue at inference time; (3) multi-prompt conditioning, which provides fine-grained and procedural control over multi-step instructions. Together, these components mitigate temporal drift, preserve motion consistency, and enable interactive video generation where future prompt updates dynamically influence ongoing streaming video generation. Experimental results demonstrate that our method produces temporally coherent and semantically faithful instructional videos that accurately follow complex, multi-step task descriptions.
title SneakPeek: Future-Guided Instructional Streaming Video Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.13019