Probing Audio-Generation Capabilities of Text-Based Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Anbazhagan, Arjun Prasaath, Kumar, Parteek, Kaur, Ujjwal, Akalin, Aslihan, Zhu, Kevin, O'Brien, Sean
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908387629858816
author Anbazhagan, Arjun Prasaath
Kumar, Parteek
Kaur, Ujjwal
Akalin, Aslihan
Zhu, Kevin
O'Brien, Sean
author_facet Anbazhagan, Arjun Prasaath
Kumar, Parteek
Kaur, Ujjwal
Akalin, Aslihan
Zhu, Kevin
O'Brien, Sean
contents How does textual representation of audio relate to the Large Language Model's (LLMs) learning about the audio world? This research investigates the extent to which LLMs can be prompted to generate audio, despite their primary training in textual data. We employ a three-tier approach, progressively increasing the complexity of audio generation: 1) Musical Notes, 2) Environmental Sounds, and 3) Human Speech. To bridge the gap between text and audio, we leverage code as an intermediary, prompting LLMs to generate code that, when executed, produces the desired audio output. To evaluate the quality and accuracy of the generated audio, we employ FAD and CLAP scores. Our findings reveal that while LLMs can generate basic audio features, their performance deteriorates as the complexity of the audio increases. This suggests that while LLMs possess a latent understanding of the auditory world, their ability to translate this understanding into tangible audio output remains rudimentary. Further research into techniques that can enhance the quality and diversity of LLM-generated audio can lead to an improvement in the performance of text-based LLMs in generating audio.
format Preprint
id arxiv_https___arxiv_org_abs_2506_00003
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Probing Audio-Generation Capabilities of Text-Based Language Models
Anbazhagan, Arjun Prasaath
Kumar, Parteek
Kaur, Ujjwal
Akalin, Aslihan
Zhu, Kevin
O'Brien, Sean
Sound
Computation and Language
Audio and Speech Processing
How does textual representation of audio relate to the Large Language Model's (LLMs) learning about the audio world? This research investigates the extent to which LLMs can be prompted to generate audio, despite their primary training in textual data. We employ a three-tier approach, progressively increasing the complexity of audio generation: 1) Musical Notes, 2) Environmental Sounds, and 3) Human Speech. To bridge the gap between text and audio, we leverage code as an intermediary, prompting LLMs to generate code that, when executed, produces the desired audio output. To evaluate the quality and accuracy of the generated audio, we employ FAD and CLAP scores. Our findings reveal that while LLMs can generate basic audio features, their performance deteriorates as the complexity of the audio increases. This suggests that while LLMs possess a latent understanding of the auditory world, their ability to translate this understanding into tangible audio output remains rudimentary. Further research into techniques that can enhance the quality and diversity of LLM-generated audio can lead to an improvement in the performance of text-based LLMs in generating audio.
title Probing Audio-Generation Capabilities of Text-Based Language Models
topic Sound
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2506.00003