The Eloquence team submission for task 1 of MLC-SLM challenge

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Concina, Lorenzo, Luque, Jordi, Brutti, Alessio, Matassoni, Marco, Zhang, Yuchen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908467268157440
author Concina, Lorenzo
Luque, Jordi
Brutti, Alessio
Matassoni, Marco
Zhang, Yuchen
author_facet Concina, Lorenzo
Luque, Jordi
Brutti, Alessio
Matassoni, Marco
Zhang, Yuchen
contents In this paper, we present our studies and experiments carried out for the task 1 of the Challenge and Workshop on Multilingual Conversational Speech Language Model (MLC-SLM), which focuses on advancing multilingual conversational speech recognition through the development of speech language models architectures. Given the increasing relevance of real-world conversational data for building robust Spoken Dialogue Systems, we explore three approaches to multilingual ASR. First, we conduct an evaluation of the official baseline to better understand its strengths and limitations, by training two projectors (linear and qformer) with different foundation models. Second we leverage the SLAM-ASR framework to train a custom multilingual linear projector. Finally we investigate the role of contrastive learning and the extended conversational context in enhancing the robustness of recognition.
format Preprint
id arxiv_https___arxiv_org_abs_2507_19308
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Eloquence team submission for task 1 of MLC-SLM challenge
Concina, Lorenzo
Luque, Jordi
Brutti, Alessio
Matassoni, Marco
Zhang, Yuchen
Sound
Computation and Language
Audio and Speech Processing
In this paper, we present our studies and experiments carried out for the task 1 of the Challenge and Workshop on Multilingual Conversational Speech Language Model (MLC-SLM), which focuses on advancing multilingual conversational speech recognition through the development of speech language models architectures. Given the increasing relevance of real-world conversational data for building robust Spoken Dialogue Systems, we explore three approaches to multilingual ASR. First, we conduct an evaluation of the official baseline to better understand its strengths and limitations, by training two projectors (linear and qformer) with different foundation models. Second we leverage the SLAM-ASR framework to train a custom multilingual linear projector. Finally we investigate the role of contrastive learning and the extended conversational context in enhancing the robustness of recognition.
title The Eloquence team submission for task 1 of MLC-SLM challenge
topic Sound
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2507.19308