MTADiffusion: Mask Text Alignment Diffusion Model for Object Inpainting

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Jun, Liu, Ting, Wu, Yihang, Qu, Xiaochao, Liu, Luoqi, Hu, Xiaolin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913918993039360
author Huang, Jun
Liu, Ting
Wu, Yihang
Qu, Xiaochao
Liu, Luoqi
Hu, Xiaolin
author_facet Huang, Jun
Liu, Ting
Wu, Yihang
Qu, Xiaochao
Liu, Luoqi
Hu, Xiaolin
contents Advancements in generative models have enabled image inpainting models to generate content within specific regions of an image based on provided prompts and masks. However, existing inpainting methods often suffer from problems such as semantic misalignment, structural distortion, and style inconsistency. In this work, we present MTADiffusion, a Mask-Text Alignment diffusion model designed for object inpainting. To enhance the semantic capabilities of the inpainting model, we introduce MTAPipeline, an automatic solution for annotating masks with detailed descriptions. Based on the MTAPipeline, we construct a new MTADataset comprising 5 million images and 25 million mask-text pairs. Furthermore, we propose a multi-task training strategy that integrates both inpainting and edge prediction tasks to improve structural stability. To promote style consistency, we present a novel inpainting style-consistency loss using a pre-trained VGG network and the Gram matrix. Comprehensive evaluations on BrushBench and EditBench demonstrate that MTADiffusion achieves state-of-the-art performance compared to other methods.
format Preprint
id arxiv_https___arxiv_org_abs_2506_23482
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MTADiffusion: Mask Text Alignment Diffusion Model for Object Inpainting
Huang, Jun
Liu, Ting
Wu, Yihang
Qu, Xiaochao
Liu, Luoqi
Hu, Xiaolin
Computer Vision and Pattern Recognition
Advancements in generative models have enabled image inpainting models to generate content within specific regions of an image based on provided prompts and masks. However, existing inpainting methods often suffer from problems such as semantic misalignment, structural distortion, and style inconsistency. In this work, we present MTADiffusion, a Mask-Text Alignment diffusion model designed for object inpainting. To enhance the semantic capabilities of the inpainting model, we introduce MTAPipeline, an automatic solution for annotating masks with detailed descriptions. Based on the MTAPipeline, we construct a new MTADataset comprising 5 million images and 25 million mask-text pairs. Furthermore, we propose a multi-task training strategy that integrates both inpainting and edge prediction tasks to improve structural stability. To promote style consistency, we present a novel inpainting style-consistency loss using a pre-trained VGG network and the Gram matrix. Comprehensive evaluations on BrushBench and EditBench demonstrate that MTADiffusion achieves state-of-the-art performance compared to other methods.
title MTADiffusion: Mask Text Alignment Diffusion Model for Object Inpainting
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.23482