PANDORA: Pixel-wise Attention Dissolution and Latent Guidance for Zero-Shot Object Removal

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Vo, Dinh-Khoi, Nguyen, Van-Loc, Nguyen, Tam V., Tran, Minh-Triet, Le, Trung-Nghia
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918415089795072
author Vo, Dinh-Khoi
Nguyen, Van-Loc
Nguyen, Tam V.
Tran, Minh-Triet
Le, Trung-Nghia
author_facet Vo, Dinh-Khoi
Nguyen, Van-Loc
Nguyen, Tam V.
Tran, Minh-Triet
Le, Trung-Nghia
contents Removing objects from natural images is challenging due to difficulty of synthesizing semantically coherent content while preserving background integrity. Existing methods often rely on fine-tuning, prompt engineering, or inference-time optimization, yet still suffer from texture inconsistency, rigid artifacts, weak foreground-background disentanglement, and poor scalability for multi-object removal. We propose a novel zero-shot object removal framework, namely PANDORA, that operates directly on pre-trained text-to-image diffusion models, requiring no fine-tuning, prompts, or optimization. We propose Pixel-wise Attention Dissolution to remove object by nullifying the most correlated attention keys for masked pixels, effectively eliminating the object from self-attention flow and allowing background context to dominate reconstruction. We further introduce Localized Attentional Disentanglement Guidance to steer denoising toward latent manifolds favorable to clean object removal. Together, these components enable precise, non-rigid, prompt-free, and scalable multi-object erasure in a single pass. Experiments demonstrate superior visual fidelity and semantic plausibility compared to state-of-the-art methods. The project page is available at https://vdkhoi20.github.io/PANDORA.
format Preprint
id arxiv_https___arxiv_org_abs_2603_27555
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle PANDORA: Pixel-wise Attention Dissolution and Latent Guidance for Zero-Shot Object Removal
Vo, Dinh-Khoi
Nguyen, Van-Loc
Nguyen, Tam V.
Tran, Minh-Triet
Le, Trung-Nghia
Computer Vision and Pattern Recognition
Removing objects from natural images is challenging due to difficulty of synthesizing semantically coherent content while preserving background integrity. Existing methods often rely on fine-tuning, prompt engineering, or inference-time optimization, yet still suffer from texture inconsistency, rigid artifacts, weak foreground-background disentanglement, and poor scalability for multi-object removal. We propose a novel zero-shot object removal framework, namely PANDORA, that operates directly on pre-trained text-to-image diffusion models, requiring no fine-tuning, prompts, or optimization. We propose Pixel-wise Attention Dissolution to remove object by nullifying the most correlated attention keys for masked pixels, effectively eliminating the object from self-attention flow and allowing background context to dominate reconstruction. We further introduce Localized Attentional Disentanglement Guidance to steer denoising toward latent manifolds favorable to clean object removal. Together, these components enable precise, non-rigid, prompt-free, and scalable multi-object erasure in a single pass. Experiments demonstrate superior visual fidelity and semantic plausibility compared to state-of-the-art methods. The project page is available at https://vdkhoi20.github.io/PANDORA.
title PANDORA: Pixel-wise Attention Dissolution and Latent Guidance for Zero-Shot Object Removal
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.27555