Enhancing Data Quality in Federated Fine-Tuning of Foundation Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Wanru, Du, Yaxin, Lane, Nicholas Donald, Chen, Siheng, Wang, Yanfeng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910357313814528
author Zhao, Wanru
Du, Yaxin
Lane, Nicholas Donald
Chen, Siheng
Wang, Yanfeng
author_facet Zhao, Wanru
Du, Yaxin
Lane, Nicholas Donald
Chen, Siheng
Wang, Yanfeng
contents In the current landscape of foundation model training, there is a significant reliance on public domain data, which is nearing exhaustion according to recent research. To further scale up, it is crucial to incorporate collaboration among multiple specialized and high-quality private domain data sources. However, the challenge of training models locally without sharing private data presents numerous obstacles in data quality control. To tackle this issue, we propose a data quality control pipeline for federated fine-tuning of foundation models. This pipeline computes scores reflecting the quality of training data and determines a global threshold for a unified standard, aiming for improved global performance. Our experiments show that the proposed quality control pipeline facilitates the effectiveness and reliability of the model training, leading to better performance.
format Preprint
id arxiv_https___arxiv_org_abs_2403_04529
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enhancing Data Quality in Federated Fine-Tuning of Foundation Models
Zhao, Wanru
Du, Yaxin
Lane, Nicholas Donald
Chen, Siheng
Wang, Yanfeng
Machine Learning
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
In the current landscape of foundation model training, there is a significant reliance on public domain data, which is nearing exhaustion according to recent research. To further scale up, it is crucial to incorporate collaboration among multiple specialized and high-quality private domain data sources. However, the challenge of training models locally without sharing private data presents numerous obstacles in data quality control. To tackle this issue, we propose a data quality control pipeline for federated fine-tuning of foundation models. This pipeline computes scores reflecting the quality of training data and determines a global threshold for a unified standard, aiming for improved global performance. Our experiments show that the proposed quality control pipeline facilitates the effectiveness and reliability of the model training, leading to better performance.
title Enhancing Data Quality in Federated Fine-Tuning of Foundation Models
topic Machine Learning
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2403.04529