Do We Need All the Synthetic Data? Targeted Image Augmentation via Diffusion Models
Fuente:
arXiv
Saved in:
| Main Authors: | Nguyen, Dang, Li, Jiping, Zheng, Jinghao, Mirzasoleiman, Baharan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Changing the Training Data Distribution to Reduce Simplicity Bias Improves In-distribution Generalization
by: Nguyen, Dang, et al.
Published: (2024)
by: Nguyen, Dang, et al.
Published: (2024)
Data Selection for Fine-tuning Vision Language Models via Cross Modal Alignment Trajectories
by: Naharas, Nilay, et al.
Published: (2025)
by: Naharas, Nilay, et al.
Published: (2025)
Data-Efficient Contrastive Language-Image Pretraining: Prioritizing Data Quality over Quantity
by: Joshi, Siddharth, et al.
Published: (2024)
by: Joshi, Siddharth, et al.
Published: (2024)
Better Safe than Sorry: Pre-training CLIP against Targeted Data Poisoning and Backdoor Attacks
by: Yang, Wenhan, et al.
Published: (2023)
by: Yang, Wenhan, et al.
Published: (2023)
Few-shot Adaptation to Distribution Shifts By Mixing Source and Target Embeddings
by: Xue, Yihao, et al.
Published: (2023)
by: Xue, Yihao, et al.
Published: (2023)
Identifying Spurious Biases Early in Training through the Lens of Simplicity Bias
by: Yang, Yu, et al.
Published: (2023)
by: Yang, Yu, et al.
Published: (2023)
MM-GEN: Enhancing Task Performance Through Targeted Multimodal Data Curation
by: Joshi, Siddharth, et al.
Published: (2025)
by: Joshi, Siddharth, et al.
Published: (2025)
Investigating the Benefits of Projection Head for Representation Learning
by: Xue, Yihao, et al.
Published: (2024)
by: Xue, Yihao, et al.
Published: (2024)
IBMA: An Imputation-Based Mixup Augmentation Using Self-Supervised Learning for Time Series Data
by: Nguyen, Dang Nha, et al.
Published: (2025)
by: Nguyen, Dang Nha, et al.
Published: (2025)
Virtually Enriched NYU Depth V2 Dataset for Monocular Depth Estimation: Do We Need Artificial Augmentation?
by: Ignatov, Dmitry, et al.
Published: (2024)
by: Ignatov, Dmitry, et al.
Published: (2024)
Representations Shape Weak-to-Strong Generalization: Theoretical Insights and Empirical Predictions
by: Xue, Yihao, et al.
Published: (2025)
by: Xue, Yihao, et al.
Published: (2025)
Do We Always Need the Simplicity Bias? Looking for Optimal Inductive Biases in the Wild
by: Teney, Damien, et al.
Published: (2025)
by: Teney, Damien, et al.
Published: (2025)
Enhancing Image Classification in Small and Unbalanced Datasets through Synthetic Data Augmentation
by: De La Fuente, Neil, et al.
Published: (2024)
by: De La Fuente, Neil, et al.
Published: (2024)
Diffusion-Based Data Augmentation for Medical Image Segmentation
by: Nazir, Maham, et al.
Published: (2025)
by: Nazir, Maham, et al.
Published: (2025)
Good Data Is All Imitation Learning Needs
by: Samadi, Amir, et al.
Published: (2024)
by: Samadi, Amir, et al.
Published: (2024)
Reliability in Semantic Segmentation: Can We Use Synthetic Data?
by: Loiseau, Thibaut, et al.
Published: (2023)
by: Loiseau, Thibaut, et al.
Published: (2023)
Utilizing Synthetic Data for Medical Vision-Language Pre-training: Bypassing the Need for Real Images
by: Liu, Che, et al.
Published: (2023)
by: Liu, Che, et al.
Published: (2023)
Improved Generation of Synthetic Imaging Data Using Feature-Aligned Diffusion
by: Nair, Lakshmi
Published: (2024)
by: Nair, Lakshmi
Published: (2024)
Synthetic Data Augmentation for Multi-Task Chinese Porcelain Classification: A Stable Diffusion Approach
by: Ling, Ziyao, et al.
Published: (2026)
by: Ling, Ziyao, et al.
Published: (2026)
Tuning the Implicit Regularizer of Masked Diffusion Language Models: Enhancing Generalization via Insights from $k$-Parity
by: Huang, Jianhao, et al.
Published: (2026)
by: Huang, Jianhao, et al.
Published: (2026)
MambaOut: Do We Really Need Mamba for Vision?
by: Yu, Weihao, et al.
Published: (2024)
by: Yu, Weihao, et al.
Published: (2024)
Do We Need Large VLMs for Spotting Soccer Actions?
by: Chakraborty, Ritabrata, et al.
Published: (2025)
by: Chakraborty, Ritabrata, et al.
Published: (2025)
Beyond Semantic Entropy: Boosting LLM Uncertainty Quantification with Pairwise Semantic Similarity
by: Nguyen, Dang, et al.
Published: (2025)
by: Nguyen, Dang, et al.
Published: (2025)
Semantic-Guided Generative Image Augmentation Method with Diffusion Models for Image Classification
by: Li, Bohan, et al.
Published: (2023)
by: Li, Bohan, et al.
Published: (2023)
Synthetic Text Generation for Training Large Language Models via Gradient Matching
by: Nguyen, Dang, et al.
Published: (2025)
by: Nguyen, Dang, et al.
Published: (2025)
Scalable Evaluation of the Realism of Synthetic Environmental Augmentations in Images
by: Ruck, Damian J., et al.
Published: (2026)
by: Ruck, Damian J., et al.
Published: (2026)
Provably Improving Generalization of Few-Shot Models with Synthetic Data
by: Nguyen, Lan-Cuong, et al.
Published: (2025)
by: Nguyen, Lan-Cuong, et al.
Published: (2025)
OCT Data is All You Need: How Vision Transformers with and without Pre-training Benefit Imaging
by: Han, Zihao, et al.
Published: (2025)
by: Han, Zihao, et al.
Published: (2025)
Synthetically Enhanced: Unveiling Synthetic Data's Potential in Medical Imaging Research
by: Khosravi, Bardia, et al.
Published: (2023)
by: Khosravi, Bardia, et al.
Published: (2023)
DoRAN: Stabilizing Weight-Decomposed Low-Rank Adaptation via Noise Injection and Auxiliary Networks
by: Diep, Nghiem T., et al.
Published: (2025)
by: Diep, Nghiem T., et al.
Published: (2025)
Diverse Image Priors for Black-box Data-free Knowledge Distillation
by: Vo, Tri-Nhan, et al.
Published: (2026)
by: Vo, Tri-Nhan, et al.
Published: (2026)
Diffusion-based Image Generation for In-distribution Data Augmentation in Surface Defect Detection
by: Capogrosso, Luigi, et al.
Published: (2024)
by: Capogrosso, Luigi, et al.
Published: (2024)
Learning Enriched Features via Selective State Spaces Model for Efficient Image Deblurring
by: Gao, Hu, et al.
Published: (2024)
by: Gao, Hu, et al.
Published: (2024)
HandCraft: Dynamic Sign Generation for Synthetic Data Augmentation
by: Rios, Gaston Gustavo, et al.
Published: (2025)
by: Rios, Gaston Gustavo, et al.
Published: (2025)
Out-of-Domain Robustness via Targeted Augmentations
by: Gao, Irena, et al.
Published: (2023)
by: Gao, Irena, et al.
Published: (2023)
Improved Multi-Task Brain Tumour Segmentation with Synthetic Data Augmentation
by: Ferreira, André, et al.
Published: (2024)
by: Ferreira, André, et al.
Published: (2024)
All in One: A Unified Synthetic Data Pipeline for Multimodal Video Understanding
by: Rahman, Tanzila, et al.
Published: (2026)
by: Rahman, Tanzila, et al.
Published: (2026)
Data Augmentation for Surgical Scene Segmentation with Anatomy-Aware Diffusion Models
by: Venkatesh, Danush Kumar, et al.
Published: (2024)
by: Venkatesh, Danush Kumar, et al.
Published: (2024)
Denoised Diffusion for Object-Focused Image Augmentation
by: Pillai, Nisha
Published: (2025)
by: Pillai, Nisha
Published: (2025)
GrootVL: Tree Topology is All You Need in State Space Model
by: Xiao, Yicheng, et al.
Published: (2024)
by: Xiao, Yicheng, et al.
Published: (2024)
Similar Items
-
Changing the Training Data Distribution to Reduce Simplicity Bias Improves In-distribution Generalization
by: Nguyen, Dang, et al.
Published: (2024) -
Data Selection for Fine-tuning Vision Language Models via Cross Modal Alignment Trajectories
by: Naharas, Nilay, et al.
Published: (2025) -
Data-Efficient Contrastive Language-Image Pretraining: Prioritizing Data Quality over Quantity
by: Joshi, Siddharth, et al.
Published: (2024) -
Better Safe than Sorry: Pre-training CLIP against Targeted Data Poisoning and Backdoor Attacks
by: Yang, Wenhan, et al.
Published: (2023) -
Few-shot Adaptation to Distribution Shifts By Mixing Source and Target Embeddings
by: Xue, Yihao, et al.
Published: (2023)