Instance-Level Generation for Representation Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Yankun, Laskar, Zakaria, Kordopatis-Zilos, Giorgos, Garcia, Noa, Tolias, Giorgos
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911201907179520
author Wu, Yankun
Laskar, Zakaria
Kordopatis-Zilos, Giorgos
Garcia, Noa
Tolias, Giorgos
author_facet Wu, Yankun
Laskar, Zakaria
Kordopatis-Zilos, Giorgos
Garcia, Noa
Tolias, Giorgos
contents Instance-level recognition (ILR) focuses on identifying individual objects rather than broad categories, offering the highest granularity in image classification. However, this fine-grained nature makes creating large-scale annotated datasets challenging, limiting ILR's real-world applicability across domains. To overcome this, we introduce a novel approach that synthetically generates diverse object instances from multiple domains under varied conditions and backgrounds, forming a large-scale training set. Unlike prior work on automatic data synthesis, our method is the first to address ILR-specific challenges without relying on any real images. Fine-tuning foundation vision models on the generated data significantly improves retrieval performance across seven ILR benchmarks spanning multiple domains. Our approach offers a new, efficient, and effective alternative to extensive data collection and curation, introducing a new ILR paradigm where the only input is the names of the target domains, unlocking a wide range of real-world applications.
format Preprint
id arxiv_https___arxiv_org_abs_2510_09171
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Instance-Level Generation for Representation Learning
Wu, Yankun
Laskar, Zakaria
Kordopatis-Zilos, Giorgos
Garcia, Noa
Tolias, Giorgos
Computer Vision and Pattern Recognition
Instance-level recognition (ILR) focuses on identifying individual objects rather than broad categories, offering the highest granularity in image classification. However, this fine-grained nature makes creating large-scale annotated datasets challenging, limiting ILR's real-world applicability across domains. To overcome this, we introduce a novel approach that synthetically generates diverse object instances from multiple domains under varied conditions and backgrounds, forming a large-scale training set. Unlike prior work on automatic data synthesis, our method is the first to address ILR-specific challenges without relying on any real images. Fine-tuning foundation vision models on the generated data significantly improves retrieval performance across seven ILR benchmarks spanning multiple domains. Our approach offers a new, efficient, and effective alternative to extensive data collection and curation, introducing a new ILR paradigm where the only input is the names of the target domains, unlocking a wide range of real-world applications.
title Instance-Level Generation for Representation Learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.09171