Late Fusion and Multi-Level Fission Amplify Cross-Modal Transfer in Text-Speech LMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cuervo, Santiago, Moumen, Adel, Labrak, Yanis, Khurana, Sameer, Laurent, Antoine, Rouvier, Mickael, Woodland, Phil, Marxer, Ricard |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Zero-Shot End-To-End Spoken Question Answering In Medical Domain
von: Labrak, Yanis, et al.
Veröffentlicht: (2024)
von: Labrak, Yanis, et al.
Veröffentlicht: (2024)
An Empirical Analysis of Discrete Unit Representations in Speech Language Modeling Pre-training
von: Labrak, Yanis, et al.
Veröffentlicht: (2025)
von: Labrak, Yanis, et al.
Veröffentlicht: (2025)
A Zero-shot and Few-shot Study of Instruction-Finetuned Large Language Models Applied to Clinical and Biomedical Tasks
von: Labrak, Yanis, et al.
Veröffentlicht: (2023)
von: Labrak, Yanis, et al.
Veröffentlicht: (2023)
Scaling Properties of Speech Language Models
von: Cuervo, Santiago, et al.
Veröffentlicht: (2024)
von: Cuervo, Santiago, et al.
Veröffentlicht: (2024)
Speech foundation models on intelligibility prediction for hearing-impaired listeners
von: Cuervo, Santiago, et al.
Veröffentlicht: (2024)
von: Cuervo, Santiago, et al.
Veröffentlicht: (2024)
How Important Is Tokenization in French Medical Masked Language Models?
von: Labrak, Yanis, et al.
Veröffentlicht: (2024)
von: Labrak, Yanis, et al.
Veröffentlicht: (2024)
BioMistral: A Collection of Open-Source Pretrained Large Language Models for Medical Domains
von: Labrak, Yanis, et al.
Veröffentlicht: (2024)
von: Labrak, Yanis, et al.
Veröffentlicht: (2024)
Transfer Learning from Whisper for Microscopic Intelligibility Prediction
von: Best, Paul, et al.
Veröffentlicht: (2024)
von: Best, Paul, et al.
Veröffentlicht: (2024)
Cross-Lingual Interleaving for Speech Language Models
von: Moumen, Adel, et al.
Veröffentlicht: (2025)
von: Moumen, Adel, et al.
Veröffentlicht: (2025)
Measuring the Redundancy of Decoder Layers in SpeechLLMs
von: Moumen, Adel, et al.
Veröffentlicht: (2026)
von: Moumen, Adel, et al.
Veröffentlicht: (2026)
Aligning Multimodal Representations through an Information Bottleneck
von: Almudévar, Antonio, et al.
Veröffentlicht: (2025)
von: Almudévar, Antonio, et al.
Veröffentlicht: (2025)
Improved Cross-Lingual Transfer Learning For Automatic Speech Translation
von: Khurana, Sameer, et al.
Veröffentlicht: (2023)
von: Khurana, Sameer, et al.
Veröffentlicht: (2023)
Probing the Information Encoded in Neural-based Acoustic Models of Automatic Speech Recognition Systems
von: Raymondaud, Quentin, et al.
Veröffentlicht: (2024)
von: Raymondaud, Quentin, et al.
Veröffentlicht: (2024)
MSP-Podcast SER Challenge 2024: L'antenne du Ventoux Multimodal Self-Supervised Learning for Speech Emotion Recognition
von: Duret, Jarod, et al.
Veröffentlicht: (2024)
von: Duret, Jarod, et al.
Veröffentlicht: (2024)
Robust Training of Vector Quantized Bottleneck Models
von: Łańcucki, Adrian, et al.
Veröffentlicht: (2020)
von: Łańcucki, Adrian, et al.
Veröffentlicht: (2020)
A Benchmark of French ASR Systems Based on Error Severity
von: Tholly, Antoine, et al.
Veröffentlicht: (2025)
von: Tholly, Antoine, et al.
Veröffentlicht: (2025)
DrBenchmark: A Large Language Understanding Evaluation Benchmark for French Biomedical Domain
von: Labrak, Yanis, et al.
Veröffentlicht: (2024)
von: Labrak, Yanis, et al.
Veröffentlicht: (2024)
Doctor or Patient? Synergizing Diarization and ASR for Code-Switched Hinglish Medical Conditions Extraction
von: Baroudi, Séverin, et al.
Veröffentlicht: (2026)
von: Baroudi, Séverin, et al.
Veröffentlicht: (2026)
Asymmetric and trial-dependent modeling: the contribution of LIA to SdSV Challenge Task 2
von: Bousquet, Pierre-Michel, et al.
Veröffentlicht: (2024)
von: Bousquet, Pierre-Michel, et al.
Veröffentlicht: (2024)
Qualitative Evaluation of Language Model Rescoring in Automatic Speech Recognition
von: Bañeras-Roux, Thibault, et al.
Veröffentlicht: (2026)
von: Bañeras-Roux, Thibault, et al.
Veröffentlicht: (2026)
A Paradigm for Interpreting Metrics and Identifying Critical Errors in Automatic Speech Recognition
von: Bañeras-Roux, Thibault, et al.
Veröffentlicht: (2026)
von: Bañeras-Roux, Thibault, et al.
Veröffentlicht: (2026)
SDialog: A Python Toolkit for End-to-End Agent Building, User Simulation, Dialog Generation, and Evaluation
von: Burdisso, Sergio, et al.
Veröffentlicht: (2025)
von: Burdisso, Sergio, et al.
Veröffentlicht: (2025)
Crossing the Species Divide: Transfer Learning from Speech to Animal Sounds
von: Cauzinille, Jules, et al.
Veröffentlicht: (2025)
von: Cauzinille, Jules, et al.
Veröffentlicht: (2025)
Factorized RVQ-GAN For Disentangled Speech Tokenization
von: Khurana, Sameer, et al.
Veröffentlicht: (2025)
von: Khurana, Sameer, et al.
Veröffentlicht: (2025)
A Comprehensive Analysis of Tokenization and Self-Supervised Learning in End-to-End Automatic Speech Recognition applied on French Language
von: Bañeras-Roux, Thibault, et al.
Veröffentlicht: (2026)
von: Bañeras-Roux, Thibault, et al.
Veröffentlicht: (2026)
Contrastive prediction strategies for unsupervised segmentation and categorization of phonemes and words
von: Cuervo, Santiago, et al.
Veröffentlicht: (2021)
von: Cuervo, Santiago, et al.
Veröffentlicht: (2021)
SDialog: A Python Toolkit for End-to-End Agent Building, User Simulation, Dialog Generation, and Evaluation
von: Burdisso, Sergio, et al.
Veröffentlicht: (2025)
von: Burdisso, Sergio, et al.
Veröffentlicht: (2025)
HATS: An Open data set Integrating Human Perception Applied to the Evaluation of Automatic Speech Recognition Metrics
von: Roux, Thibault Bañeras, et al.
Veröffentlicht: (2026)
von: Roux, Thibault Bañeras, et al.
Veröffentlicht: (2026)
MT2KD: Towards A General-Purpose Encoder for Speech, Speaker, and Audio Events
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2024)
PixIT: Joint Training of Speaker Diarization and Speech Separation from Real-world Multi-speaker Recordings
von: Kalda, Joonas, et al.
Veröffentlicht: (2024)
von: Kalda, Joonas, et al.
Veröffentlicht: (2024)
Entropy-Aligned Decoding of LMs for Better Writing and Reasoning
von: Ahmed, Kareem, et al.
Veröffentlicht: (2026)
von: Ahmed, Kareem, et al.
Veröffentlicht: (2026)
Synthetic Lyrics Detection Across Languages and Genres
von: Labrak, Yanis, et al.
Veröffentlicht: (2024)
von: Labrak, Yanis, et al.
Veröffentlicht: (2024)
Depth Jitter: Seeing through the Depth
von: Rahman, Md Sazidur, et al.
Veröffentlicht: (2025)
von: Rahman, Md Sazidur, et al.
Veröffentlicht: (2025)
Generating Synthetic Doctor-Patient Conversations for Long-form Audio Summarization
von: Labrak, Yanis, et al.
Veröffentlicht: (2026)
von: Labrak, Yanis, et al.
Veröffentlicht: (2026)
ProGRes: Prompted Generative Rescoring on ASR n-Best
von: Tur, Ada Defne, et al.
Veröffentlicht: (2024)
von: Tur, Ada Defne, et al.
Veröffentlicht: (2024)
Discrete Audio Tokens: More Than a Survey!
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2025)
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2025)
Enhancing Multi-Corpus Training in SSL-Based Anti-Spoofing Models: Domain-Invariant Feature Extraction
von: Dao, Anh-Tuan, et al.
Veröffentlicht: (2026)
von: Dao, Anh-Tuan, et al.
Veröffentlicht: (2026)
On the Use of Self-Supervised Representation Learning for Speaker Diarization and Separation
von: Baroudi, Séverin, et al.
Veröffentlicht: (2025)
von: Baroudi, Séverin, et al.
Veröffentlicht: (2025)
Deep Learning Classification With Noisy Labels
von: Sanchez, Guillaume, et al.
Veröffentlicht: (2020)
von: Sanchez, Guillaume, et al.
Veröffentlicht: (2020)
Applying machine learning to primate bioacoustics: Review and perspectives
von: Jules Cauzinille, et al.
Veröffentlicht: (2024)
von: Jules Cauzinille, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Zero-Shot End-To-End Spoken Question Answering In Medical Domain
von: Labrak, Yanis, et al.
Veröffentlicht: (2024) -
An Empirical Analysis of Discrete Unit Representations in Speech Language Modeling Pre-training
von: Labrak, Yanis, et al.
Veröffentlicht: (2025) -
A Zero-shot and Few-shot Study of Instruction-Finetuned Large Language Models Applied to Clinical and Biomedical Tasks
von: Labrak, Yanis, et al.
Veröffentlicht: (2023) -
Scaling Properties of Speech Language Models
von: Cuervo, Santiago, et al.
Veröffentlicht: (2024) -
Speech foundation models on intelligibility prediction for hearing-impaired listeners
von: Cuervo, Santiago, et al.
Veröffentlicht: (2024)