Boosting Few-Shot Detection with Large Language Models and Layout-to-Image Synthesis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Abdullah, Ahmed, Ebert, Nikolas, Wasenmüller, Oliver
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916429270351872
author Abdullah, Ahmed
Ebert, Nikolas
Wasenmüller, Oliver
author_facet Abdullah, Ahmed
Ebert, Nikolas
Wasenmüller, Oliver
contents Recent advancements in diffusion models have enabled a wide range of works exploiting their ability to generate high-volume, high-quality data for use in various downstream tasks. One subclass of such models, dubbed Layout-to-Image Synthesis (LIS), learns to generate images conditioned on a spatial layout (bounding boxes, masks, poses, etc.) and has shown a promising ability to generate realistic images, albeit with limited layout-adherence. Moreover, the question of how to effectively transfer those models for scalable augmentation of few-shot detection data remains unanswered. Thus, we propose a collaborative framework employing a Large Language Model (LLM) and an LIS model for enhancing few-shot detection beyond state-of-the-art generative augmentation approaches. We leverage LLM's reasoning ability to extrapolate the spatial prior of the annotation space by generating new bounding boxes given only a few example annotations. Additionally, we introduce our novel layout-aware CLIP score for sample ranking, enabling tight coupling between generated layouts and images. Significant improvements on COCO few-shot benchmarks are observed. With our approach, a YOLOX-S baseline is boosted by more than 140%, 50%, 35% in mAP on the COCO 5-,10-, and 30-shot settings, respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2410_06841
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Boosting Few-Shot Detection with Large Language Models and Layout-to-Image Synthesis
Abdullah, Ahmed
Ebert, Nikolas
Wasenmüller, Oliver
Computer Vision and Pattern Recognition
Recent advancements in diffusion models have enabled a wide range of works exploiting their ability to generate high-volume, high-quality data for use in various downstream tasks. One subclass of such models, dubbed Layout-to-Image Synthesis (LIS), learns to generate images conditioned on a spatial layout (bounding boxes, masks, poses, etc.) and has shown a promising ability to generate realistic images, albeit with limited layout-adherence. Moreover, the question of how to effectively transfer those models for scalable augmentation of few-shot detection data remains unanswered. Thus, we propose a collaborative framework employing a Large Language Model (LLM) and an LIS model for enhancing few-shot detection beyond state-of-the-art generative augmentation approaches. We leverage LLM's reasoning ability to extrapolate the spatial prior of the annotation space by generating new bounding boxes given only a few example annotations. Additionally, we introduce our novel layout-aware CLIP score for sample ranking, enabling tight coupling between generated layouts and images. Significant improvements on COCO few-shot benchmarks are observed. With our approach, a YOLOX-S baseline is boosted by more than 140%, 50%, 35% in mAP on the COCO 5-,10-, and 30-shot settings, respectively.
title Boosting Few-Shot Detection with Large Language Models and Layout-to-Image Synthesis
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.06841