Investigating Cost-Efficiency of LLM-Generated Training Data for Conversational Semantic Frame Analysis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Matta, Shiho, Huang, Yin Jou, Cheng, Fei, Kiyomaru, Hirokazu, Murawaki, Yugo
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912065010008064
author Matta, Shiho
Huang, Yin Jou
Cheng, Fei
Kiyomaru, Hirokazu
Murawaki, Yugo
author_facet Matta, Shiho
Huang, Yin Jou
Cheng, Fei
Kiyomaru, Hirokazu
Murawaki, Yugo
contents Recent studies have demonstrated that few-shot learning allows LLMs to generate training data for supervised models at a low cost. However, the quality of LLM-generated data may not entirely match that of human-labeled data. This raises a crucial question: how should one balance the trade-off between the higher quality but more expensive human data and the lower quality yet substantially cheaper LLM-generated data? In this paper, we synthesized training data for conversational semantic frame analysis using GPT-4 and examined how to allocate budgets optimally to achieve the best performance. Our experiments, conducted across various budget levels, reveal that optimal cost-efficiency is achieved by combining both human and LLM-generated data across a wide range of budget levels. Notably, as the budget decreases, a higher proportion of LLM-generated data becomes more preferable.
format Preprint
id arxiv_https___arxiv_org_abs_2410_06550
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Investigating Cost-Efficiency of LLM-Generated Training Data for Conversational Semantic Frame Analysis
Matta, Shiho
Huang, Yin Jou
Cheng, Fei
Kiyomaru, Hirokazu
Murawaki, Yugo
Computation and Language
Artificial Intelligence
Recent studies have demonstrated that few-shot learning allows LLMs to generate training data for supervised models at a low cost. However, the quality of LLM-generated data may not entirely match that of human-labeled data. This raises a crucial question: how should one balance the trade-off between the higher quality but more expensive human data and the lower quality yet substantially cheaper LLM-generated data? In this paper, we synthesized training data for conversational semantic frame analysis using GPT-4 and examined how to allocate budgets optimally to achieve the best performance. Our experiments, conducted across various budget levels, reveal that optimal cost-efficiency is achieved by combining both human and LLM-generated data across a wide range of budget levels. Notably, as the budget decreases, a higher proportion of LLM-generated data becomes more preferable.
title Investigating Cost-Efficiency of LLM-Generated Training Data for Conversational Semantic Frame Analysis
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2410.06550