SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866911575744446464 |
|---|---|
| author | Xiao, Yicheng Zhang, Wenhu Song, Lin Chen, Yukang Li, Wenbo Jiang, Nan Ren, Tianhe Lin, Haokun Huang, Wei Huang, Haoyang Li, Xiu Duan, Nan Qi, Xiaojuan |
| author_facet | Xiao, Yicheng Zhang, Wenhu Song, Lin Chen, Yukang Li, Wenbo Jiang, Nan Ren, Tianhe Lin, Haokun Huang, Wei Huang, Haoyang Li, Xiu Duan, Nan Qi, Xiaojuan |
| contents | Image spatial editing performs geometry-driven transformations, allowing precise control over object layout and camera viewpoints. Current models are insufficient for fine-grained spatial manipulations, motivating a dedicated assessment suite. Our contributions are listed: (i) We introduce SpatialEdit-Bench, a complete benchmark that evaluates spatial editing by jointly measuring perceptual plausibility and geometric fidelity via viewpoint reconstruction and framing analysis. (ii) To address the data bottleneck for scalable training, we construct SpatialEdit-500k, a synthetic dataset generated with a controllable Blender pipeline that renders objects across diverse backgrounds and systematic camera trajectories, providing precise ground-truth transformations for both object- and camera-centric operations. (iii) Building on this data, we develop SpatialEdit-16B, a baseline model for fine-grained spatial editing. Our method achieves competitive performance on general editing while substantially outperforming prior methods on spatial manipulation tasks. All resources will be made public at https://github.com/EasonXiao-888/SpatialEdit. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_04911 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing Xiao, Yicheng Zhang, Wenhu Song, Lin Chen, Yukang Li, Wenbo Jiang, Nan Ren, Tianhe Lin, Haokun Huang, Wei Huang, Haoyang Li, Xiu Duan, Nan Qi, Xiaojuan Computer Vision and Pattern Recognition Image spatial editing performs geometry-driven transformations, allowing precise control over object layout and camera viewpoints. Current models are insufficient for fine-grained spatial manipulations, motivating a dedicated assessment suite. Our contributions are listed: (i) We introduce SpatialEdit-Bench, a complete benchmark that evaluates spatial editing by jointly measuring perceptual plausibility and geometric fidelity via viewpoint reconstruction and framing analysis. (ii) To address the data bottleneck for scalable training, we construct SpatialEdit-500k, a synthetic dataset generated with a controllable Blender pipeline that renders objects across diverse backgrounds and systematic camera trajectories, providing precise ground-truth transformations for both object- and camera-centric operations. (iii) Building on this data, we develop SpatialEdit-16B, a baseline model for fine-grained spatial editing. Our method achieves competitive performance on general editing while substantially outperforming prior methods on spatial manipulation tasks. All resources will be made public at https://github.com/EasonXiao-888/SpatialEdit. |
| title | SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2604.04911 |