Is Synthetic Data all We Need? Benchmarking the Robustness of Models Trained with Synthetic Images

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Singh, Krishnakant, Navaratnam, Thanush, Holmer, Jannik, Schaub-Meyer, Simone, Roth, Stefan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913408883884032
author Singh, Krishnakant
Navaratnam, Thanush
Holmer, Jannik
Schaub-Meyer, Simone
Roth, Stefan
author_facet Singh, Krishnakant
Navaratnam, Thanush
Holmer, Jannik
Schaub-Meyer, Simone
Roth, Stefan
contents A long-standing challenge in developing machine learning approaches has been the lack of high-quality labeled data. Recently, models trained with purely synthetic data, here termed synthetic clones, generated using large-scale pre-trained diffusion models have shown promising results in overcoming this annotation bottleneck. As these synthetic clone models progress, they are likely to be deployed in challenging real-world settings, yet their suitability remains understudied. Our work addresses this gap by providing the first benchmark for three classes of synthetic clone models, namely supervised, self-supervised, and multi-modal ones, across a range of robustness measures. We show that existing synthetic self-supervised and multi-modal clones are comparable to or outperform state-of-the-art real-image baselines for a range of robustness metrics - shape bias, background bias, calibration, etc. However, we also find that synthetic clones are much more susceptible to adversarial and real-world noise than models trained with real data. To address this, we find that combining both real and synthetic data further increases the robustness, and that the choice of prompt used for generating synthetic images plays an important part in the robustness of synthetic clones.
format Preprint
id arxiv_https___arxiv_org_abs_2405_20469
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Is Synthetic Data all We Need? Benchmarking the Robustness of Models Trained with Synthetic Images
Singh, Krishnakant
Navaratnam, Thanush
Holmer, Jannik
Schaub-Meyer, Simone
Roth, Stefan
Computer Vision and Pattern Recognition
A long-standing challenge in developing machine learning approaches has been the lack of high-quality labeled data. Recently, models trained with purely synthetic data, here termed synthetic clones, generated using large-scale pre-trained diffusion models have shown promising results in overcoming this annotation bottleneck. As these synthetic clone models progress, they are likely to be deployed in challenging real-world settings, yet their suitability remains understudied. Our work addresses this gap by providing the first benchmark for three classes of synthetic clone models, namely supervised, self-supervised, and multi-modal ones, across a range of robustness measures. We show that existing synthetic self-supervised and multi-modal clones are comparable to or outperform state-of-the-art real-image baselines for a range of robustness metrics - shape bias, background bias, calibration, etc. However, we also find that synthetic clones are much more susceptible to adversarial and real-world noise than models trained with real data. To address this, we find that combining both real and synthetic data further increases the robustness, and that the choice of prompt used for generating synthetic images plays an important part in the robustness of synthetic clones.
title Is Synthetic Data all We Need? Benchmarking the Robustness of Models Trained with Synthetic Images
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.20469