Training Flow Matching Models with Reliable Labels via Self-Purification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Hyeongju, Yu, Yechan, Yi, June Young, Lee, Juheon
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916964647043072
author Kim, Hyeongju
Yu, Yechan
Yi, June Young
Lee, Juheon
author_facet Kim, Hyeongju
Yu, Yechan
Yi, June Young
Lee, Juheon
contents Training datasets are inherently imperfect, often containing mislabeled samples due to human annotation errors, limitations of tagging models, and other sources of noise. Such label contamination can significantly degrade the performance of a trained model. In this work, we introduce Self-Purifying Flow Matching (SPFM), a principled approach to filtering unreliable data within the flow-matching framework. SPFM identifies suspicious data using the model itself during the training process, bypassing the need for pretrained models or additional modules. Our experiments demonstrate that models trained with SPFM generate samples that accurately adhere to the specified conditioning, even when trained on noisy labels. Furthermore, we validate the robustness of SPFM on the TITW dataset, which consists of in-the-wild speech data, achieving performance that surpasses existing baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2509_19091
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Training Flow Matching Models with Reliable Labels via Self-Purification
Kim, Hyeongju
Yu, Yechan
Yi, June Young
Lee, Juheon
Audio and Speech Processing
Artificial Intelligence
Sound
Training datasets are inherently imperfect, often containing mislabeled samples due to human annotation errors, limitations of tagging models, and other sources of noise. Such label contamination can significantly degrade the performance of a trained model. In this work, we introduce Self-Purifying Flow Matching (SPFM), a principled approach to filtering unreliable data within the flow-matching framework. SPFM identifies suspicious data using the model itself during the training process, bypassing the need for pretrained models or additional modules. Our experiments demonstrate that models trained with SPFM generate samples that accurately adhere to the specified conditioning, even when trained on noisy labels. Furthermore, we validate the robustness of SPFM on the TITW dataset, which consists of in-the-wild speech data, achieving performance that surpasses existing baselines.
title Training Flow Matching Models with Reliable Labels via Self-Purification
topic Audio and Speech Processing
Artificial Intelligence
Sound
url https://arxiv.org/abs/2509.19091