Image Augmentation Agent for Weakly Supervised Semantic Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Wangyu, Qiu, Xianglin, Song, Siqi, Chen, Zhenhong, Huang, Xiaowei, Ma, Fei, Xiao, Jimin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912550437781504
author Wu, Wangyu
Qiu, Xianglin
Song, Siqi
Chen, Zhenhong
Huang, Xiaowei
Ma, Fei
Xiao, Jimin
author_facet Wu, Wangyu
Qiu, Xianglin
Song, Siqi
Chen, Zhenhong
Huang, Xiaowei
Ma, Fei
Xiao, Jimin
contents Weakly-supervised semantic segmentation (WSSS) has achieved remarkable progress using only image-level labels. However, most existing WSSS methods focus on designing new network structures and loss functions to generate more accurate dense labels, overlooking the limitations imposed by fixed datasets, which can constrain performance improvements. We argue that more diverse trainable images provides WSSS richer information and help model understand more comprehensive semantic pattern. Therefore in this paper, we introduce a novel approach called Image Augmentation Agent (IAA) which shows that it is possible to enhance WSSS from data generation perspective. IAA mainly design an augmentation agent that leverages large language models (LLMs) and diffusion models to automatically generate additional images for WSSS. In practice, to address the instability in prompt generation by LLMs, we develop a prompt self-refinement mechanism. It allow LLMs to re-evaluate the rationality of generated prompts to produce more coherent prompts. Additionally, we insert an online filter into diffusion generation process to dynamically ensure the quality and balance of generated images. Experimental results show that our method significantly surpasses state-of-the-art WSSS approaches on the PASCAL VOC 2012 and MS COCO 2014 datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2412_20439
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Image Augmentation Agent for Weakly Supervised Semantic Segmentation
Wu, Wangyu
Qiu, Xianglin
Song, Siqi
Chen, Zhenhong
Huang, Xiaowei
Ma, Fei
Xiao, Jimin
Computer Vision and Pattern Recognition
Weakly-supervised semantic segmentation (WSSS) has achieved remarkable progress using only image-level labels. However, most existing WSSS methods focus on designing new network structures and loss functions to generate more accurate dense labels, overlooking the limitations imposed by fixed datasets, which can constrain performance improvements. We argue that more diverse trainable images provides WSSS richer information and help model understand more comprehensive semantic pattern. Therefore in this paper, we introduce a novel approach called Image Augmentation Agent (IAA) which shows that it is possible to enhance WSSS from data generation perspective. IAA mainly design an augmentation agent that leverages large language models (LLMs) and diffusion models to automatically generate additional images for WSSS. In practice, to address the instability in prompt generation by LLMs, we develop a prompt self-refinement mechanism. It allow LLMs to re-evaluate the rationality of generated prompts to produce more coherent prompts. Additionally, we insert an online filter into diffusion generation process to dynamically ensure the quality and balance of generated images. Experimental results show that our method significantly surpasses state-of-the-art WSSS approaches on the PASCAL VOC 2012 and MS COCO 2014 datasets.
title Image Augmentation Agent for Weakly Supervised Semantic Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.20439