Mitigating Long-Tail Bias via Prompt-Controlled Diffusion Augmentation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wijenayake, Buddhi, Wasalathilake, Nichula, Godaliyadda, Roshan, Herath, Vijitha, Ekanayake, Parakrama, Patel, Vishal M.
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908992793477120
author Wijenayake, Buddhi
Wasalathilake, Nichula
Godaliyadda, Roshan
Herath, Vijitha
Ekanayake, Parakrama
Patel, Vishal M.
author_facet Wijenayake, Buddhi
Wasalathilake, Nichula
Godaliyadda, Roshan
Herath, Vijitha
Ekanayake, Parakrama
Patel, Vishal M.
contents Long-tailed class imbalance remains a fundamental obstacle in semantic segmentation of high-resolution remote-sensing imagery, where dominant classes shape learned representations and rare classes are systematically under-segmented. This challenge becomes more acute in cross-domain settings such as LoveDA, which exhibits an explicit Urban/Rural split with substantial appearance differences and inconsistent class-frequency statistics across domains. We propose a prompt-controlled diffusion augmentation framework that generates paired label-image samples with explicit control over semantic composition and domain, enabling targeted enrichment of underrepresented classes rather than indiscriminate dataset expansion. A domain-aware, masked, ratio-conditioned discrete diffusion model first synthesizes layouts that satisfy class-ratio targets while preserving realistic spatial co-occurrence, and a ControlNet-guided diffusion model then renders photorealistic, domain-consistent images from these layouts. When mixed with real data, the resulting synthetic pairs improve multiple segmentation backbones, especially on minority classes and under domain shift, showing that better downstream segmentation comes from adding the right samples in the right proportions. Source codes, pretrained models, and synthetic datasets are available at \href{https://buddhi19.github.io/SyntheticGen}{\texttt{buddhi19.github.io/SyntheticGen}}.
format Preprint
id arxiv_https___arxiv_org_abs_2602_04749
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Mitigating Long-Tail Bias via Prompt-Controlled Diffusion Augmentation
Wijenayake, Buddhi
Wasalathilake, Nichula
Godaliyadda, Roshan
Herath, Vijitha
Ekanayake, Parakrama
Patel, Vishal M.
Computer Vision and Pattern Recognition
Long-tailed class imbalance remains a fundamental obstacle in semantic segmentation of high-resolution remote-sensing imagery, where dominant classes shape learned representations and rare classes are systematically under-segmented. This challenge becomes more acute in cross-domain settings such as LoveDA, which exhibits an explicit Urban/Rural split with substantial appearance differences and inconsistent class-frequency statistics across domains. We propose a prompt-controlled diffusion augmentation framework that generates paired label-image samples with explicit control over semantic composition and domain, enabling targeted enrichment of underrepresented classes rather than indiscriminate dataset expansion. A domain-aware, masked, ratio-conditioned discrete diffusion model first synthesizes layouts that satisfy class-ratio targets while preserving realistic spatial co-occurrence, and a ControlNet-guided diffusion model then renders photorealistic, domain-consistent images from these layouts. When mixed with real data, the resulting synthetic pairs improve multiple segmentation backbones, especially on minority classes and under domain shift, showing that better downstream segmentation comes from adding the right samples in the right proportions. Source codes, pretrained models, and synthetic datasets are available at \href{https://buddhi19.github.io/SyntheticGen}{\texttt{buddhi19.github.io/SyntheticGen}}.
title Mitigating Long-Tail Bias via Prompt-Controlled Diffusion Augmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.04749