The Unmet Promise of Synthetic Training Images: Using Retrieved Real Images Performs Better
Fuente:
arXiv
Saved in:
| Main Authors: | Geng, Scott, Hsieh, Cheng-Yu, Ramanujan, Vivek, Wallingford, Matthew, Li, Chun-Liang, Koh, Pang Wei, Krishna, Ranjay |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Posterior Augmented Flow Matching
by: Stoica, George, et al.
Published: (2026)
by: Stoica, George, et al.
Published: (2026)
The Delta Learning Hypothesis: Preference Tuning on Weak Data can Yield Strong Gains
by: Geng, Scott, et al.
Published: (2025)
by: Geng, Scott, et al.
Published: (2025)
Multilingual Diversity Improves Vision-Language Representations
by: Nguyen, Thao, et al.
Published: (2024)
by: Nguyen, Thao, et al.
Published: (2024)
Contrastive Flow Matching
by: Stoica, George, et al.
Published: (2025)
by: Stoica, George, et al.
Published: (2025)
From an Image to a Scene: Learning to Imagine the World from a Million 360 Videos
by: Wallingford, Matthew, et al.
Published: (2024)
by: Wallingford, Matthew, et al.
Published: (2024)
FocalLens: Instruction Tuning Enables Zero-Shot Conditional Image Representations
by: Hsieh, Cheng-Yu, et al.
Published: (2025)
by: Hsieh, Cheng-Yu, et al.
Published: (2025)
Negative Token Merging: Image-based Adversarial Feature Guidance
by: Singh, Jaskirat, et al.
Published: (2024)
by: Singh, Jaskirat, et al.
Published: (2024)
The Hard Positive Truth about Vision-Language Compositionality
by: Kamath, Amita, et al.
Published: (2024)
by: Kamath, Amita, et al.
Published: (2024)
When Worse is Better: Navigating the compression-generation tradeoff in visual tokenization
by: Ramanujan, Vivek, et al.
Published: (2024)
by: Ramanujan, Vivek, et al.
Published: (2024)
EVE: Enabling Anyone to Train Robots using Augmented Reality
by: Wang, Jun, et al.
Published: (2024)
by: Wang, Jun, et al.
Published: (2024)
OmniView: An All-Seeing Diffusion Model for 3D and 4D View Synthesis
by: Fan, Xiang, et al.
Published: (2025)
by: Fan, Xiang, et al.
Published: (2025)
VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition
by: Yadav, Tanush, et al.
Published: (2026)
by: Yadav, Tanush, et al.
Published: (2026)
Spurious Rewards: Rethinking Training Signals in RLVR
by: Shao, Rulin, et al.
Published: (2025)
by: Shao, Rulin, et al.
Published: (2025)
Agonistic Image Generation: Unsettling the Hegemony of Intention
by: Shaw, Andrew, et al.
Published: (2025)
by: Shaw, Andrew, et al.
Published: (2025)
Lookback Lens: Detecting and Mitigating Contextual Hallucinations in Large Language Models Using Only Attention Maps
by: Chuang, Yung-Sung, et al.
Published: (2024)
by: Chuang, Yung-Sung, et al.
Published: (2024)
REALEDIT: Reddit Edits As a Large-scale Empirical Dataset for Image Transformations
by: Sushko, Peter, et al.
Published: (2025)
by: Sushko, Peter, et al.
Published: (2025)
Can Synthetic Query Rewrites Capture User Intent Better than Humans in Retrieval-Augmented Generation?
by: Zheng, JiaYing, et al.
Published: (2025)
by: Zheng, JiaYing, et al.
Published: (2025)
Matryoshka Representation Learning
by: Kusupati, Aditya, et al.
Published: (2022)
by: Kusupati, Aditya, et al.
Published: (2022)
Mask-aware Text-to-Image Retrieval: Referring Expression Segmentation Meets Cross-modal Retrieval
by: Shen, Li-Cheng, et al.
Published: (2025)
by: Shen, Li-Cheng, et al.
Published: (2025)
Ablate-to-Validate: Are Vision-Language Models Really Using Continuous Thought Tokens?
by: Zhang, Tianyi, et al.
Published: (2026)
by: Zhang, Tianyi, et al.
Published: (2026)
Polyp Detection Using YOLOv9 on Real and Synthetic Colonoscopy Images
by: Seo Yi Chng, et al.
Published: (2026)
by: Seo Yi Chng, et al.
Published: (2026)
Improving Text Generation on Images with Synthetic Captions
by: Koh, Jun Young, et al.
Published: (2024)
by: Koh, Jun Young, et al.
Published: (2024)
GenEval 2: Addressing Benchmark Drift in Text-to-Image Evaluation
by: Kamath, Amita, et al.
Published: (2025)
by: Kamath, Amita, et al.
Published: (2025)
Synthetic Visual Genome
by: Park, Jae Sung, et al.
Published: (2025)
by: Park, Jae Sung, et al.
Published: (2025)
Invisible Relevance Bias: Text-Image Retrieval Models Prefer AI-Generated Images
by: Xu, Shicheng, et al.
Published: (2023)
by: Xu, Shicheng, et al.
Published: (2023)
JPEG-LM: LLMs as Image Generators with Canonical Codec Representations
by: Han, Xiaochuang, et al.
Published: (2024)
by: Han, Xiaochuang, et al.
Published: (2024)
XR: Cross-Modal Agents for Composed Image Retrieval
by: Yang, Zhongyu, et al.
Published: (2026)
by: Yang, Zhongyu, et al.
Published: (2026)
Semantic and Expressive Variation in Image Captions Across Languages
by: Ye, Andre, et al.
Published: (2023)
by: Ye, Andre, et al.
Published: (2023)
Algebraic generators of the skein algebra of a surface
by: Santharoubane, Ramanujan
Published: (2018)
by: Santharoubane, Ramanujan
Published: (2018)
Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data Generation
by: Yang, Yue, et al.
Published: (2025)
by: Yang, Yue, et al.
Published: (2025)
Superposed Decoding: Multiple Generations from a Single Autoregressive Inference Pass
by: Shen, Ethan, et al.
Published: (2024)
by: Shen, Ethan, et al.
Published: (2024)
Language-only Efficient Training of Zero-shot Composed Image Retrieval
by: Gu, Geonmo, et al.
Published: (2023)
by: Gu, Geonmo, et al.
Published: (2023)
Synthetic Data Augmentation for Table Detection: Re-evaluating TableNet's Performance with Automatically Generated Document Images
by: Sahukara, Krishna, et al.
Published: (2025)
by: Sahukara, Krishna, et al.
Published: (2025)
Better Synthetic Data by Retrieving and Transforming Existing Datasets
by: Gandhi, Saumya, et al.
Published: (2024)
by: Gandhi, Saumya, et al.
Published: (2024)
When Rubrics Fail: Error Enumeration as Reward in Reference-Free RL Post-Training for Virtual Try-On
by: Ikezogwo, Wisdom, et al.
Published: (2026)
by: Ikezogwo, Wisdom, et al.
Published: (2026)
Diffusion Curriculum: Synthetic-to-Real Data Curriculum via Image-Guided Diffusion
by: Liang, Yijun, et al.
Published: (2024)
by: Liang, Yijun, et al.
Published: (2024)
MIMIC: Masked Image Modeling with Image Correspondences
by: Marathe, Kalyani, et al.
Published: (2023)
by: Marathe, Kalyani, et al.
Published: (2023)
Image-Text Out-Of-Context Detection Using Synthetic Multimodal Misinformation
by: Shalabi, Fatma, et al.
Published: (2024)
by: Shalabi, Fatma, et al.
Published: (2024)
Generate Any Scene: Scene Graph Driven Data Synthesis for Visual Generation Training
by: Gao, Ziqi, et al.
Published: (2024)
by: Gao, Ziqi, et al.
Published: (2024)
Graph-Based Captioning: Enhancing Visual Descriptions by Interconnecting Region Captions
by: Hsieh, Yu-Guan, et al.
Published: (2024)
by: Hsieh, Yu-Guan, et al.
Published: (2024)
Similar Items
-
Posterior Augmented Flow Matching
by: Stoica, George, et al.
Published: (2026) -
The Delta Learning Hypothesis: Preference Tuning on Weak Data can Yield Strong Gains
by: Geng, Scott, et al.
Published: (2025) -
Multilingual Diversity Improves Vision-Language Representations
by: Nguyen, Thao, et al.
Published: (2024) -
Contrastive Flow Matching
by: Stoica, George, et al.
Published: (2025) -
From an Image to a Scene: Learning to Imagine the World from a Million 360 Videos
by: Wallingford, Matthew, et al.
Published: (2024)