Preservation of Language Understanding Capabilities in Speech-aware Large Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kubis, Marek, Skórzewski, Paweł, Christop, Iwona, Czyżnikiewicz, Mateusz, Kubiak, Jakub, Bondaruk, Łukasz, Lewandowski, Marcin
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918161477009408
author Kubis, Marek
Skórzewski, Paweł
Christop, Iwona
Czyżnikiewicz, Mateusz
Kubiak, Jakub
Bondaruk, Łukasz
Lewandowski, Marcin
author_facet Kubis, Marek
Skórzewski, Paweł
Christop, Iwona
Czyżnikiewicz, Mateusz
Kubiak, Jakub
Bondaruk, Łukasz
Lewandowski, Marcin
contents The paper presents C3T (Cross-modal Capabilities Conservation Test), a new benchmark for assessing the performance of speech-aware large language models. The benchmark utilizes textual tasks and a voice cloning text-to-speech model to quantify the extent to which language understanding capabilities are preserved when the model is accessed via speech input. C3T quantifies the fairness of the model for different categories of speakers and its robustness across text and speech modalities.
format Preprint
id arxiv_https___arxiv_org_abs_2509_12171
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Preservation of Language Understanding Capabilities in Speech-aware Large Language Models
Kubis, Marek
Skórzewski, Paweł
Christop, Iwona
Czyżnikiewicz, Mateusz
Kubiak, Jakub
Bondaruk, Łukasz
Lewandowski, Marcin
Computation and Language
Artificial Intelligence
The paper presents C3T (Cross-modal Capabilities Conservation Test), a new benchmark for assessing the performance of speech-aware large language models. The benchmark utilizes textual tasks and a voice cloning text-to-speech model to quantify the extent to which language understanding capabilities are preserved when the model is accessed via speech input. C3T quantifies the fairness of the model for different categories of speakers and its robustness across text and speech modalities.
title Preservation of Language Understanding Capabilities in Speech-aware Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2509.12171