Alchemist: Unlocking Efficiency in Text-to-Image Model Training via Meta-Gradient Data Selection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ding, Kaixin, Zhou, Yang, Chen, Xi, Yang, Miao, Ou, Jiarong, Chen, Rui, Tao, Xin, Zhao, Hengshuang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914208298303488
author Ding, Kaixin
Zhou, Yang
Chen, Xi
Yang, Miao
Ou, Jiarong
Chen, Rui
Tao, Xin
Zhao, Hengshuang
author_facet Ding, Kaixin
Zhou, Yang
Chen, Xi
Yang, Miao
Ou, Jiarong
Chen, Rui
Tao, Xin
Zhao, Hengshuang
contents Recent advances in Text-to-Image (T2I) generative models, such as Imagen, Stable Diffusion, and FLUX, have led to remarkable improvements in visual quality. However, their performance is fundamentally limited by the quality of training data. Web-crawled and synthetic image datasets often contain low-quality or redundant samples, which lead to degraded visual fidelity, unstable training, and inefficient computation. Hence, effective data selection is crucial for improving data efficiency. Existing approaches rely on costly manual curation or heuristic scoring based on single-dimensional features in Text-to-Image data filtering. Although meta-learning based method has been explored in LLM, there is no adaptation for image modalities. To this end, we propose **Alchemist**, a meta-gradient-based framework to select a suitable subset from large-scale text-image data pairs. Our approach automatically learns to assess the influence of each sample by iteratively optimizing the model from a data-centric perspective. Alchemist consists of two key stages: data rating and data pruning. We train a lightweight rater to estimate each sample's influence based on gradient information, enhanced with multi-granularity perception. We then use the Shift-Gsampling strategy to select informative subsets for efficient model training. Alchemist is the first automatic, scalable, meta-gradient-based data selection framework for Text-to-Image model training. Experiments on both synthetic and web-crawled datasets demonstrate that Alchemist consistently improves visual quality and downstream performance. Training on an Alchemist-selected 50% of the data can outperform training on the full dataset.
format Preprint
id arxiv_https___arxiv_org_abs_2512_16905
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Alchemist: Unlocking Efficiency in Text-to-Image Model Training via Meta-Gradient Data Selection
Ding, Kaixin
Zhou, Yang
Chen, Xi
Yang, Miao
Ou, Jiarong
Chen, Rui
Tao, Xin
Zhao, Hengshuang
Computer Vision and Pattern Recognition
Recent advances in Text-to-Image (T2I) generative models, such as Imagen, Stable Diffusion, and FLUX, have led to remarkable improvements in visual quality. However, their performance is fundamentally limited by the quality of training data. Web-crawled and synthetic image datasets often contain low-quality or redundant samples, which lead to degraded visual fidelity, unstable training, and inefficient computation. Hence, effective data selection is crucial for improving data efficiency. Existing approaches rely on costly manual curation or heuristic scoring based on single-dimensional features in Text-to-Image data filtering. Although meta-learning based method has been explored in LLM, there is no adaptation for image modalities. To this end, we propose **Alchemist**, a meta-gradient-based framework to select a suitable subset from large-scale text-image data pairs. Our approach automatically learns to assess the influence of each sample by iteratively optimizing the model from a data-centric perspective. Alchemist consists of two key stages: data rating and data pruning. We train a lightweight rater to estimate each sample's influence based on gradient information, enhanced with multi-granularity perception. We then use the Shift-Gsampling strategy to select informative subsets for efficient model training. Alchemist is the first automatic, scalable, meta-gradient-based data selection framework for Text-to-Image model training. Experiments on both synthetic and web-crawled datasets demonstrate that Alchemist consistently improves visual quality and downstream performance. Training on an Alchemist-selected 50% of the data can outperform training on the full dataset.
title Alchemist: Unlocking Efficiency in Text-to-Image Model Training via Meta-Gradient Data Selection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.16905