Preserve, Reveal, Expand: Faithful 4D Video Editing with Region-Aware Conditioning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Hu, Zhangchi, Sun, Wenzhang, Yin, Xiangchen, Yuan, Jiahui, Wang, Chunfeng, Li, Hao, Zhan, Kun, Sun, Xiaoyan
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913148087304192
author Hu, Zhangchi
Sun, Wenzhang
Yin, Xiangchen
Yuan, Jiahui
Wang, Chunfeng
Li, Hao
Zhan, Kun
Sun, Xiaoyan
author_facet Hu, Zhangchi
Sun, Wenzhang
Yin, Xiangchen
Yuan, Jiahui
Wang, Chunfeng
Li, Hao
Zhan, Kun
Sun, Xiaoyan
contents Existing 4D-driven video diffusion models primarily target plausible generation, but faithful 4D editing requires preserving source-observed regions while synthesizing disoccluded or out-of-view content. We identify Evidence-Role Mismatch: reliable source-backed evidence, unreliable rendered cues, and unsupported regions are entangled in a single conditioning signal, causing preservation drift, ghosting, and unstable extrapolation. We propose PREX (Preserve, Reveal, Expand), a region-aware framework that decomposes the target spatiotemporal volume into Preserve, Reveal, and Expand roles according to observation support and scene extent. PREX builds observation-backed appearance cues with calibrated confidence and injects them into a frozen video diffusion backbone through a region-aware adapter, trained with proxy tasks without requiring paired edited videos. We further introduce PREBench, a diagnostic benchmark with curated edits, region-role masks, and human-aligned metrics that complement global video-quality and 4D-control evaluations. Experiments show that PREX reduces region-structured failures while maintaining strong visual quality and 4D edit control capability. Project Page: https://ricepastem.github.io/PREX-Open
format Preprint
id arxiv_https___arxiv_org_abs_2605_20961
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Preserve, Reveal, Expand: Faithful 4D Video Editing with Region-Aware Conditioning
Hu, Zhangchi
Sun, Wenzhang
Yin, Xiangchen
Yuan, Jiahui
Wang, Chunfeng
Li, Hao
Zhan, Kun
Sun, Xiaoyan
Computer Vision and Pattern Recognition
Existing 4D-driven video diffusion models primarily target plausible generation, but faithful 4D editing requires preserving source-observed regions while synthesizing disoccluded or out-of-view content. We identify Evidence-Role Mismatch: reliable source-backed evidence, unreliable rendered cues, and unsupported regions are entangled in a single conditioning signal, causing preservation drift, ghosting, and unstable extrapolation. We propose PREX (Preserve, Reveal, Expand), a region-aware framework that decomposes the target spatiotemporal volume into Preserve, Reveal, and Expand roles according to observation support and scene extent. PREX builds observation-backed appearance cues with calibrated confidence and injects them into a frozen video diffusion backbone through a region-aware adapter, trained with proxy tasks without requiring paired edited videos. We further introduce PREBench, a diagnostic benchmark with curated edits, region-role masks, and human-aligned metrics that complement global video-quality and 4D-control evaluations. Experiments show that PREX reduces region-structured failures while maintaining strong visual quality and 4D edit control capability. Project Page: https://ricepastem.github.io/PREX-Open
title Preserve, Reveal, Expand: Faithful 4D Video Editing with Region-Aware Conditioning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.20961