QUOTA: Quantifying Objects with Text-to-Image Models for Any Domain

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Sun, Wenfang, Du, Yingjun, Liu, Gaowen, Zheng, Yefeng, Snoek, Cees G. M.
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917145481314304
author Sun, Wenfang
Du, Yingjun
Liu, Gaowen
Zheng, Yefeng
Snoek, Cees G. M.
author_facet Sun, Wenfang
Du, Yingjun
Liu, Gaowen
Zheng, Yefeng
Snoek, Cees G. M.
contents We tackle the problem of quantifying the number of objects by a generative text-to-image model. Rather than retraining such a model for each new image domain of interest, which leads to high computational costs and limited scalability, we are the first to consider this problem from a domain-agnostic perspective. We propose QUOTA, an optimization framework for text-to-image models that enables effective object quantification across unseen domains without retraining. It leverages a dual-loop meta-learning strategy to optimize a domain-invariant prompt. Further, by integrating prompt learning with learnable counting and domain tokens, our method captures stylistic variations and maintains accuracy, even for object classes not encountered during training. For evaluation, we adopt a new benchmark specifically designed for object quantification in domain generalization, enabling rigorous assessment of object quantification accuracy and adaptability across unseen domains in text-to-image generation. Extensive experiments demonstrate that QUOTA outperforms conventional models in both object quantification accuracy and semantic consistency, setting a new benchmark for efficient and scalable text-to-image generation for any domain.
format Preprint
id arxiv_https___arxiv_org_abs_2411_19534
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle QUOTA: Quantifying Objects with Text-to-Image Models for Any Domain
Sun, Wenfang
Du, Yingjun
Liu, Gaowen
Zheng, Yefeng
Snoek, Cees G. M.
Computer Vision and Pattern Recognition
Machine Learning
We tackle the problem of quantifying the number of objects by a generative text-to-image model. Rather than retraining such a model for each new image domain of interest, which leads to high computational costs and limited scalability, we are the first to consider this problem from a domain-agnostic perspective. We propose QUOTA, an optimization framework for text-to-image models that enables effective object quantification across unseen domains without retraining. It leverages a dual-loop meta-learning strategy to optimize a domain-invariant prompt. Further, by integrating prompt learning with learnable counting and domain tokens, our method captures stylistic variations and maintains accuracy, even for object classes not encountered during training. For evaluation, we adopt a new benchmark specifically designed for object quantification in domain generalization, enabling rigorous assessment of object quantification accuracy and adaptability across unseen domains in text-to-image generation. Extensive experiments demonstrate that QUOTA outperforms conventional models in both object quantification accuracy and semantic consistency, setting a new benchmark for efficient and scalable text-to-image generation for any domain.
title QUOTA: Quantifying Objects with Text-to-Image Models for Any Domain
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2411.19534