SOEDiff: Efficient Distillation for Small Object Editing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Yiming, Pan, Qihe, Zhao, Zhen, Wang, Zicheng, Long, Sifan, Liang, Ronghua
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916546291433472
author Wu, Yiming
Pan, Qihe
Zhao, Zhen
Wang, Zicheng
Long, Sifan
Liang, Ronghua
author_facet Wu, Yiming
Pan, Qihe
Zhao, Zhen
Wang, Zicheng
Long, Sifan
Liang, Ronghua
contents In this paper, we delve into a new task known as small object editing (SOE), which focuses on text-based image inpainting within a constrained, small-sized area. Despite the remarkable success have been achieved by current image inpainting approaches, their application to the SOE task generally results in failure cases such as Object Missing, Text-Image Mismatch, and Distortion. These failures stem from the limited use of small-sized objects in training datasets and the downsampling operations employed by U-Net models, which hinders accurate generation. To overcome these challenges, we introduce a novel training-based approach, SOEDiff, aimed at enhancing the capability of baseline models like StableDiffusion in editing small-sized objects while minimizing training costs. Specifically, our method involves two key components: SO-LoRA, which efficiently fine-tunes low-rank matrices, and Cross-Scale Score Distillation loss, which leverages high-resolution predictions from the pre-trained teacher diffusion model. Our method presents significant improvements on the test dataset collected from MSCOCO and OpenImage, validating the effectiveness of our proposed method in small object editing. In particular, when comparing SOEDiff with SD-I model on the OpenImage-f dataset, we observe a 0.99 improvement in CLIP-Score and a reduction of 2.87 in FID.
format Preprint
id arxiv_https___arxiv_org_abs_2405_09114
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SOEDiff: Efficient Distillation for Small Object Editing
Wu, Yiming
Pan, Qihe
Zhao, Zhen
Wang, Zicheng
Long, Sifan
Liang, Ronghua
Computer Vision and Pattern Recognition
In this paper, we delve into a new task known as small object editing (SOE), which focuses on text-based image inpainting within a constrained, small-sized area. Despite the remarkable success have been achieved by current image inpainting approaches, their application to the SOE task generally results in failure cases such as Object Missing, Text-Image Mismatch, and Distortion. These failures stem from the limited use of small-sized objects in training datasets and the downsampling operations employed by U-Net models, which hinders accurate generation. To overcome these challenges, we introduce a novel training-based approach, SOEDiff, aimed at enhancing the capability of baseline models like StableDiffusion in editing small-sized objects while minimizing training costs. Specifically, our method involves two key components: SO-LoRA, which efficiently fine-tunes low-rank matrices, and Cross-Scale Score Distillation loss, which leverages high-resolution predictions from the pre-trained teacher diffusion model. Our method presents significant improvements on the test dataset collected from MSCOCO and OpenImage, validating the effectiveness of our proposed method in small object editing. In particular, when comparing SOEDiff with SD-I model on the OpenImage-f dataset, we observe a 0.99 improvement in CLIP-Score and a reduction of 2.87 in FID.
title SOEDiff: Efficient Distillation for Small Object Editing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.09114