Fill In The Gaps: Model Calibration and Generalization with Synthetic Data

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ba, Yang, Mancenido, Michelle V., Pan, Rong
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912072615329792
author Ba, Yang
Mancenido, Michelle V.
Pan, Rong
author_facet Ba, Yang
Mancenido, Michelle V.
Pan, Rong
contents As machine learning models continue to swiftly advance, calibrating their performance has become a major concern prior to practical and widespread implementation. Most existing calibration methods often negatively impact model accuracy due to the lack of diversity of validation data, resulting in reduced generalizability. To address this, we propose a calibration method that incorporates synthetic data without compromising accuracy. We derive the expected calibration error (ECE) bound using the Probably Approximately Correct (PAC) learning framework. Large language models (LLMs), known for their ability to mimic real data and generate text with mixed class labels, are utilized as a synthetic data generation strategy to lower the ECE bound and improve model accuracy on real test data. Additionally, we propose data generation mechanisms for efficient calibration. Testing our method on four different natural language processing tasks, we observed an average up to 34\% increase in accuracy and 33\% decrease in ECE.
format Preprint
id arxiv_https___arxiv_org_abs_2410_10864
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Fill In The Gaps: Model Calibration and Generalization with Synthetic Data
Ba, Yang
Mancenido, Michelle V.
Pan, Rong
Computation and Language
Artificial Intelligence
Machine Learning
As machine learning models continue to swiftly advance, calibrating their performance has become a major concern prior to practical and widespread implementation. Most existing calibration methods often negatively impact model accuracy due to the lack of diversity of validation data, resulting in reduced generalizability. To address this, we propose a calibration method that incorporates synthetic data without compromising accuracy. We derive the expected calibration error (ECE) bound using the Probably Approximately Correct (PAC) learning framework. Large language models (LLMs), known for their ability to mimic real data and generate text with mixed class labels, are utilized as a synthetic data generation strategy to lower the ECE bound and improve model accuracy on real test data. Additionally, we propose data generation mechanisms for efficient calibration. Testing our method on four different natural language processing tasks, we observed an average up to 34\% increase in accuracy and 33\% decrease in ECE.
title Fill In The Gaps: Model Calibration and Generalization with Synthetic Data
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2410.10864