VOX-KRIKRI: Unifying Speech and Language through Continuous Fusion

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Damianos, Dimitrios, Voukoutis, Leon, Paraskevopoulos, Georgios, Katsouros, Vassilis
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918144248905728
author Damianos, Dimitrios
Voukoutis, Leon
Paraskevopoulos, Georgios
Katsouros, Vassilis
author_facet Damianos, Dimitrios
Voukoutis, Leon
Paraskevopoulos, Georgios
Katsouros, Vassilis
contents We present a multimodal fusion framework that bridges pre-trained decoder-based large language models (LLM) and acoustic encoder-decoder architectures such as Whisper, with the aim of building speech-enabled LLMs. Instead of directly using audio embeddings, we explore an intermediate audio-conditioned text space as a more effective mechanism for alignment. Our method operates fully in continuous text representation spaces, fusing Whisper's hidden decoder states with those of an LLM through cross-modal attention, and supports both offline and streaming modes. We introduce \textit{VoxKrikri}, the first Greek speech LLM, and show through analysis that our approach effectively aligns representations across modalities. These results highlight continuous space fusion as a promising path for multilingual and low-resource speech LLMs, while achieving state-of-the-art results for Automatic Speech Recognition in Greek, providing an average $\sim20\%$ relative improvement across benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2509_15667
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VOX-KRIKRI: Unifying Speech and Language through Continuous Fusion
Damianos, Dimitrios
Voukoutis, Leon
Paraskevopoulos, Georgios
Katsouros, Vassilis
Computation and Language
Sound
Audio and Speech Processing
We present a multimodal fusion framework that bridges pre-trained decoder-based large language models (LLM) and acoustic encoder-decoder architectures such as Whisper, with the aim of building speech-enabled LLMs. Instead of directly using audio embeddings, we explore an intermediate audio-conditioned text space as a more effective mechanism for alignment. Our method operates fully in continuous text representation spaces, fusing Whisper's hidden decoder states with those of an LLM through cross-modal attention, and supports both offline and streaming modes. We introduce \textit{VoxKrikri}, the first Greek speech LLM, and show through analysis that our approach effectively aligns representations across modalities. These results highlight continuous space fusion as a promising path for multilingual and low-resource speech LLMs, while achieving state-of-the-art results for Automatic Speech Recognition in Greek, providing an average $\sim20\%$ relative improvement across benchmarks.
title VOX-KRIKRI: Unifying Speech and Language through Continuous Fusion
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2509.15667