GlobalPaint: Spatiotemporal Coherent Video Outpainting with Global Feature Guidance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pan, Yueming, Feng, Ruoyu, Bao, Jianmin, Luo, Chong, Zheng, Nanning
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915720067022848
author Pan, Yueming
Feng, Ruoyu
Bao, Jianmin
Luo, Chong
Zheng, Nanning
author_facet Pan, Yueming
Feng, Ruoyu
Bao, Jianmin
Luo, Chong
Zheng, Nanning
contents Video outpainting extends a video beyond its original boundaries by synthesizing missing border content. Compared with image outpainting, it requires not only per-frame spatial plausibility but also long-range temporal coherence, especially when outpainted content becomes visible across time under camera or object motion. We propose GlobalPaint, a diffusion-based framework for spatiotemporal coherent video outpainting. Our approach adopts a hierarchical pipeline that first outpaints key frames and then completes intermediate frames via an interpolation model conditioned on the completed boundaries, reducing error accumulation in sequential processing. At the model level, we augment a pretrained image inpainting backbone with (i) an Enhanced Spatial-Temporal module featuring 3D windowed attention for stronger spatiotemporal interaction, and (ii) global feature guidance that distills OpenCLIP features from observed regions across all frames into compact global tokens using a dedicated extractor. Comprehensive evaluations on benchmark datasets demonstrate improved reconstruction quality and more natural motion compared to prior methods. Our demo page is https://yuemingpan.github.io/GlobalPaint/
format Preprint
id arxiv_https___arxiv_org_abs_2601_06413
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle GlobalPaint: Spatiotemporal Coherent Video Outpainting with Global Feature Guidance
Pan, Yueming
Feng, Ruoyu
Bao, Jianmin
Luo, Chong
Zheng, Nanning
Computer Vision and Pattern Recognition
Video outpainting extends a video beyond its original boundaries by synthesizing missing border content. Compared with image outpainting, it requires not only per-frame spatial plausibility but also long-range temporal coherence, especially when outpainted content becomes visible across time under camera or object motion. We propose GlobalPaint, a diffusion-based framework for spatiotemporal coherent video outpainting. Our approach adopts a hierarchical pipeline that first outpaints key frames and then completes intermediate frames via an interpolation model conditioned on the completed boundaries, reducing error accumulation in sequential processing. At the model level, we augment a pretrained image inpainting backbone with (i) an Enhanced Spatial-Temporal module featuring 3D windowed attention for stronger spatiotemporal interaction, and (ii) global feature guidance that distills OpenCLIP features from observed regions across all frames into compact global tokens using a dedicated extractor. Comprehensive evaluations on benchmark datasets demonstrate improved reconstruction quality and more natural motion compared to prior methods. Our demo page is https://yuemingpan.github.io/GlobalPaint/
title GlobalPaint: Spatiotemporal Coherent Video Outpainting with Global Feature Guidance
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.06413