From Words to Waves: Analyzing Concept Formation in Speech and Text-Based Foundation Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ersoy, Asım, Mousi, Basel, Chowdhury, Shammur, Alam, Firoj, Dalvi, Fahim, Durrani, Nadir
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915317826977792
author Ersoy, Asım
Mousi, Basel
Chowdhury, Shammur
Alam, Firoj
Dalvi, Fahim
Durrani, Nadir
author_facet Ersoy, Asım
Mousi, Basel
Chowdhury, Shammur
Alam, Firoj
Dalvi, Fahim
Durrani, Nadir
contents The emergence of large language models (LLMs) has demonstrated that systems trained solely on text can acquire extensive world knowledge, develop reasoning capabilities, and internalize abstract semantic concepts--showcasing properties that can be associated with general intelligence. This raises an intriguing question: Do such concepts emerge in models trained on other modalities, such as speech? Furthermore, when models are trained jointly on multiple modalities: Do they develop a richer, more structured semantic understanding? To explore this, we analyze the conceptual structures learned by speech and textual models both individually and jointly. We employ Latent Concept Analysis, an unsupervised method for uncovering and interpreting latent representations in neural networks, to examine how semantic abstractions form across modalities. For reproducibility we made scripts and other resources available to the community.
format Preprint
id arxiv_https___arxiv_org_abs_2506_01133
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle From Words to Waves: Analyzing Concept Formation in Speech and Text-Based Foundation Models
Ersoy, Asım
Mousi, Basel
Chowdhury, Shammur
Alam, Firoj
Dalvi, Fahim
Durrani, Nadir
Computation and Language
Artificial Intelligence
Sound
Audio and Speech Processing
The emergence of large language models (LLMs) has demonstrated that systems trained solely on text can acquire extensive world knowledge, develop reasoning capabilities, and internalize abstract semantic concepts--showcasing properties that can be associated with general intelligence. This raises an intriguing question: Do such concepts emerge in models trained on other modalities, such as speech? Furthermore, when models are trained jointly on multiple modalities: Do they develop a richer, more structured semantic understanding? To explore this, we analyze the conceptual structures learned by speech and textual models both individually and jointly. We employ Latent Concept Analysis, an unsupervised method for uncovering and interpreting latent representations in neural networks, to examine how semantic abstractions form across modalities. For reproducibility we made scripts and other resources available to the community.
title From Words to Waves: Analyzing Concept Formation in Speech and Text-Based Foundation Models
topic Computation and Language
Artificial Intelligence
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2506.01133