Training-Free Image Editing with Visual Context Integration and Concept Alignment
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866915917743521792 |
|---|---|
| author | Song, Rui Wang, Guo-Hua Chen, Qing-Guo Luo, Weihua Xu, Tongda Liu, Zhening Wang, Yan Lin, Zehong Zhang, Jun |
| author_facet | Song, Rui Wang, Guo-Hua Chen, Qing-Guo Luo, Weihua Xu, Tongda Liu, Zhening Wang, Yan Lin, Zehong Zhang, Jun |
| contents | In image editing, it is essential to incorporate a context image to convey the user's precise requirements, such as subject appearance or image style. Existing training-based visual context-aware editing methods incur data collection effort and training cost. On the other hand, the training-free alternatives are typically established on diffusion inversion, which struggles with consistency and flexibility. In this work, we propose VicoEdit, a training-free and inversion-free method to inject the visual context into the pretrained text-prompted editing model. More specifically, VicoEdit directly transforms the source image into the target one based on the visual context, thereby eliminating the need for inversion that can lead to deviated trajectories. Moreover, we design a posterior sampling approach guided by concept alignment to enhance the editing consistency. Empirical results demonstrate that our training-free method achieves even better editing performance than the state-of-the-art training-based models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_04487 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Training-Free Image Editing with Visual Context Integration and Concept Alignment Song, Rui Wang, Guo-Hua Chen, Qing-Guo Luo, Weihua Xu, Tongda Liu, Zhening Wang, Yan Lin, Zehong Zhang, Jun Computer Vision and Pattern Recognition In image editing, it is essential to incorporate a context image to convey the user's precise requirements, such as subject appearance or image style. Existing training-based visual context-aware editing methods incur data collection effort and training cost. On the other hand, the training-free alternatives are typically established on diffusion inversion, which struggles with consistency and flexibility. In this work, we propose VicoEdit, a training-free and inversion-free method to inject the visual context into the pretrained text-prompted editing model. More specifically, VicoEdit directly transforms the source image into the target one based on the visual context, thereby eliminating the need for inversion that can lead to deviated trajectories. Moreover, we design a posterior sampling approach guided by concept alignment to enhance the editing consistency. Empirical results demonstrate that our training-free method achieves even better editing performance than the state-of-the-art training-based models. |
| title | Training-Free Image Editing with Visual Context Integration and Concept Alignment |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2604.04487 |