Investigating Cost-Efficiency of LLM-Generated Training Data for Conversational Semantic Frame Analysis
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912065010008064 |
|---|---|
| author | Matta, Shiho Huang, Yin Jou Cheng, Fei Kiyomaru, Hirokazu Murawaki, Yugo |
| author_facet | Matta, Shiho Huang, Yin Jou Cheng, Fei Kiyomaru, Hirokazu Murawaki, Yugo |
| contents | Recent studies have demonstrated that few-shot learning allows LLMs to generate training data for supervised models at a low cost. However, the quality of LLM-generated data may not entirely match that of human-labeled data. This raises a crucial question: how should one balance the trade-off between the higher quality but more expensive human data and the lower quality yet substantially cheaper LLM-generated data? In this paper, we synthesized training data for conversational semantic frame analysis using GPT-4 and examined how to allocate budgets optimally to achieve the best performance. Our experiments, conducted across various budget levels, reveal that optimal cost-efficiency is achieved by combining both human and LLM-generated data across a wide range of budget levels. Notably, as the budget decreases, a higher proportion of LLM-generated data becomes more preferable. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_06550 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Investigating Cost-Efficiency of LLM-Generated Training Data for Conversational Semantic Frame Analysis Matta, Shiho Huang, Yin Jou Cheng, Fei Kiyomaru, Hirokazu Murawaki, Yugo Computation and Language Artificial Intelligence Recent studies have demonstrated that few-shot learning allows LLMs to generate training data for supervised models at a low cost. However, the quality of LLM-generated data may not entirely match that of human-labeled data. This raises a crucial question: how should one balance the trade-off between the higher quality but more expensive human data and the lower quality yet substantially cheaper LLM-generated data? In this paper, we synthesized training data for conversational semantic frame analysis using GPT-4 and examined how to allocate budgets optimally to achieve the best performance. Our experiments, conducted across various budget levels, reveal that optimal cost-efficiency is achieved by combining both human and LLM-generated data across a wide range of budget levels. Notably, as the budget decreases, a higher proportion of LLM-generated data becomes more preferable. |
| title | Investigating Cost-Efficiency of LLM-Generated Training Data for Conversational Semantic Frame Analysis |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2410.06550 |