Large Language Model Data Generation for Enhanced Intent Recognition in German Speech

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rosin, Theresa Pekarek, Kaplan, Burak Can, Wermter, Stefan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916888004526080
author Rosin, Theresa Pekarek
Kaplan, Burak Can
Wermter, Stefan
author_facet Rosin, Theresa Pekarek
Kaplan, Burak Can
Wermter, Stefan
contents Intent recognition (IR) for speech commands is essential for artificial intelligence (AI) assistant systems; however, most existing approaches are limited to short commands and are predominantly developed for English. This paper addresses these limitations by focusing on IR from speech by elderly German speakers. We propose a novel approach that combines an adapted Whisper ASR model, fine-tuned on elderly German speech (SVC-de), with Transformer-based language models trained on synthetic text datasets generated by three well-known large language models (LLMs): LeoLM, Llama3, and ChatGPT. To evaluate the robustness of our approach, we generate synthetic speech with a text-to-speech model and conduct extensive cross-dataset testing. Our results show that synthetic LLM-generated data significantly boosts classification performance and robustness to different speaking styles and unseen vocabulary. Notably, we find that LeoLM, a smaller, domain-specific 13B LLM, surpasses the much larger ChatGPT (175B) in dataset quality for German intent recognition. Our approach demonstrates that generative AI can effectively bridge data gaps in low-resource domains. We provide detailed documentation of our data generation and training process to ensure transparency and reproducibility.
format Preprint
id arxiv_https___arxiv_org_abs_2508_06277
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Large Language Model Data Generation for Enhanced Intent Recognition in German Speech
Rosin, Theresa Pekarek
Kaplan, Burak Can
Wermter, Stefan
Computation and Language
Machine Learning
Sound
Intent recognition (IR) for speech commands is essential for artificial intelligence (AI) assistant systems; however, most existing approaches are limited to short commands and are predominantly developed for English. This paper addresses these limitations by focusing on IR from speech by elderly German speakers. We propose a novel approach that combines an adapted Whisper ASR model, fine-tuned on elderly German speech (SVC-de), with Transformer-based language models trained on synthetic text datasets generated by three well-known large language models (LLMs): LeoLM, Llama3, and ChatGPT. To evaluate the robustness of our approach, we generate synthetic speech with a text-to-speech model and conduct extensive cross-dataset testing. Our results show that synthetic LLM-generated data significantly boosts classification performance and robustness to different speaking styles and unseen vocabulary. Notably, we find that LeoLM, a smaller, domain-specific 13B LLM, surpasses the much larger ChatGPT (175B) in dataset quality for German intent recognition. Our approach demonstrates that generative AI can effectively bridge data gaps in low-resource domains. We provide detailed documentation of our data generation and training process to ensure transparency and reproducibility.
title Large Language Model Data Generation for Enhanced Intent Recognition in German Speech
topic Computation and Language
Machine Learning
Sound
url https://arxiv.org/abs/2508.06277