The Unmet Promise of Synthetic Training Images: Using Retrieved Real Images Performs Better

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Geng, Scott, Hsieh, Cheng-Yu, Ramanujan, Vivek, Wallingford, Matthew, Li, Chun-Liang, Koh, Pang Wei, Krishna, Ranjay
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917881767264256
author Geng, Scott
Hsieh, Cheng-Yu
Ramanujan, Vivek
Wallingford, Matthew
Li, Chun-Liang
Koh, Pang Wei
Krishna, Ranjay
author_facet Geng, Scott
Hsieh, Cheng-Yu
Ramanujan, Vivek
Wallingford, Matthew
Li, Chun-Liang
Koh, Pang Wei
Krishna, Ranjay
contents Generative text-to-image models enable us to synthesize unlimited amounts of images in a controllable manner, spurring many recent efforts to train vision models with synthetic data. However, every synthetic image ultimately originates from the upstream data used to train the generator. Does the intermediate generator provide additional information over directly training on relevant parts of the upstream data? Grounding this question in the setting of image classification, we compare finetuning on task-relevant, targeted synthetic data generated by Stable Diffusion -- a generative model trained on the LAION-2B dataset -- against finetuning on targeted real images retrieved directly from LAION-2B. We show that while synthetic data can benefit some downstream tasks, it is universally matched or outperformed by real data from the simple retrieval baseline. Our analysis suggests that this underperformance is partially due to generator artifacts and inaccurate task-relevant visual details in the synthetic images. Overall, we argue that targeted retrieval is a critical baseline to consider when training with synthetic data -- a baseline that current methods do not yet surpass. We release code, data, and models at https://github.com/scottgeng00/unmet-promise.
format Preprint
id arxiv_https___arxiv_org_abs_2406_05184
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle The Unmet Promise of Synthetic Training Images: Using Retrieved Real Images Performs Better
Geng, Scott
Hsieh, Cheng-Yu
Ramanujan, Vivek
Wallingford, Matthew
Li, Chun-Liang
Koh, Pang Wei
Krishna, Ranjay
Computer Vision and Pattern Recognition
Generative text-to-image models enable us to synthesize unlimited amounts of images in a controllable manner, spurring many recent efforts to train vision models with synthetic data. However, every synthetic image ultimately originates from the upstream data used to train the generator. Does the intermediate generator provide additional information over directly training on relevant parts of the upstream data? Grounding this question in the setting of image classification, we compare finetuning on task-relevant, targeted synthetic data generated by Stable Diffusion -- a generative model trained on the LAION-2B dataset -- against finetuning on targeted real images retrieved directly from LAION-2B. We show that while synthetic data can benefit some downstream tasks, it is universally matched or outperformed by real data from the simple retrieval baseline. Our analysis suggests that this underperformance is partially due to generator artifacts and inaccurate task-relevant visual details in the synthetic images. Overall, we argue that targeted retrieval is a critical baseline to consider when training with synthetic data -- a baseline that current methods do not yet surpass. We release code, data, and models at https://github.com/scottgeng00/unmet-promise.
title The Unmet Promise of Synthetic Training Images: Using Retrieved Real Images Performs Better
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2406.05184