Edit3r: Instant 3D Scene Editing from Sparse Unposed Images

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Jiageng, Lyu, Weijie, Li, Xueting, Guo, Yejie, Yang, Ming-Hsuan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917178585907200
author Liu, Jiageng
Lyu, Weijie
Li, Xueting
Guo, Yejie
Yang, Ming-Hsuan
author_facet Liu, Jiageng
Lyu, Weijie
Li, Xueting
Guo, Yejie
Yang, Ming-Hsuan
contents We present Edit3r, a feed-forward framework that reconstructs and edits 3D scenes in a single pass from unposed, view-inconsistent, instruction-edited images. Unlike prior methods requiring per-scene optimization, Edit3r directly predicts instruction-aligned 3D edits, enabling fast and photorealistic rendering without optimization or pose estimation. A key challenge in training such a model lies in the absence of multi-view consistent edited images for supervision. We address this with (i) a SAM2-based recoloring strategy that generates reliable, cross-view-consistent supervision, and (ii) an asymmetric input strategy that pairs a recolored reference view with raw auxiliary views, encouraging the network to fuse and align disparate observations. At inference, our model effectively handles images edited by 2D methods such as InstructPix2Pix, despite not being exposed to such edits during training. For large-scale quantitative evaluation, we introduce DL3DV-Edit-Bench, a benchmark built on the DL3DV test split, featuring 20 diverse scenes, 4 edit types and 100 edits in total. Comprehensive quantitative and qualitative results show that Edit3r achieves superior semantic alignment and enhanced 3D consistency compared to recent baselines, while operating at significantly higher inference speed, making it promising for real-time 3D editing applications.
format Preprint
id arxiv_https___arxiv_org_abs_2512_25071
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Edit3r: Instant 3D Scene Editing from Sparse Unposed Images
Liu, Jiageng
Lyu, Weijie
Li, Xueting
Guo, Yejie
Yang, Ming-Hsuan
Computer Vision and Pattern Recognition
We present Edit3r, a feed-forward framework that reconstructs and edits 3D scenes in a single pass from unposed, view-inconsistent, instruction-edited images. Unlike prior methods requiring per-scene optimization, Edit3r directly predicts instruction-aligned 3D edits, enabling fast and photorealistic rendering without optimization or pose estimation. A key challenge in training such a model lies in the absence of multi-view consistent edited images for supervision. We address this with (i) a SAM2-based recoloring strategy that generates reliable, cross-view-consistent supervision, and (ii) an asymmetric input strategy that pairs a recolored reference view with raw auxiliary views, encouraging the network to fuse and align disparate observations. At inference, our model effectively handles images edited by 2D methods such as InstructPix2Pix, despite not being exposed to such edits during training. For large-scale quantitative evaluation, we introduce DL3DV-Edit-Bench, a benchmark built on the DL3DV test split, featuring 20 diverse scenes, 4 edit types and 100 edits in total. Comprehensive quantitative and qualitative results show that Edit3r achieves superior semantic alignment and enhanced 3D consistency compared to recent baselines, while operating at significantly higher inference speed, making it promising for real-time 3D editing applications.
title Edit3r: Instant 3D Scene Editing from Sparse Unposed Images
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.25071