mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition
Fuente:
arXiv
Guardado en:
| Autores principales: | Rouditchenko, Andrew, Thomas, Samuel, Kuehne, Hilde, Feris, Rogerio, Glass, James |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation
por: Rouditchenko, Andrew, et al.
Publicado: (2024)
por: Rouditchenko, Andrew, et al.
Publicado: (2024)
Omni-R1: Do You Really Need Audio to Fine-Tune Your Audio LLM?
por: Rouditchenko, Andrew, et al.
Publicado: (2025)
por: Rouditchenko, Andrew, et al.
Publicado: (2025)
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment
por: Araujo, Edson, et al.
Publicado: (2025)
por: Araujo, Edson, et al.
Publicado: (2025)
Towards Audio Token Compression in Large Audio Language Models
por: Bhati, Saurabhchand, et al.
Publicado: (2025)
por: Bhati, Saurabhchand, et al.
Publicado: (2025)
DASS: Distilled Audio State Space Models Are Stronger and More Duration-Scalable Learners
por: Bhati, Saurabhchand, et al.
Publicado: (2024)
por: Bhati, Saurabhchand, et al.
Publicado: (2024)
State-Space Large Audio Language Models
por: Bhati, Saurabhchand, et al.
Publicado: (2024)
por: Bhati, Saurabhchand, et al.
Publicado: (2024)
Uncovering the Visual Contribution in Audio-Visual Speech Recognition
por: Lin, Zhaofeng, et al.
Publicado: (2024)
por: Lin, Zhaofeng, et al.
Publicado: (2024)
DQ-Whisper: Joint Distillation and Quantization for Efficient Multilingual Speech Recognition
por: Shao, Hang, et al.
Publicado: (2023)
por: Shao, Hang, et al.
Publicado: (2023)
Sparsely Shared LoRA on Whisper for Child Speech Recognition
por: Liu, Wei, et al.
Publicado: (2023)
por: Liu, Wei, et al.
Publicado: (2023)
Leveraging Self-Supervised Models for Automatic Whispered Speech Recognition
por: Farhadipour, Aref, et al.
Publicado: (2024)
por: Farhadipour, Aref, et al.
Publicado: (2024)
A Study on Incorporating Whisper for Robust Speech Assessment
por: Zezario, Ryandhimas E., et al.
Publicado: (2023)
por: Zezario, Ryandhimas E., et al.
Publicado: (2023)
Human-Inspired Computing for Robust and Efficient Audio-Visual Speech Recognition
por: Liu, Qianhui, et al.
Publicado: (2024)
por: Liu, Qianhui, et al.
Publicado: (2024)
Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding
por: Zhou, Haoran, et al.
Publicado: (2025)
por: Zhou, Haoran, et al.
Publicado: (2025)
State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition
por: Farhadipour, Aref, et al.
Publicado: (2025)
por: Farhadipour, Aref, et al.
Publicado: (2025)
A Self-Training Approach for Whisper to Enhance Long Dysarthric Speech Recognition
por: Wang, Shiyao, et al.
Publicado: (2025)
por: Wang, Shiyao, et al.
Publicado: (2025)
Noisy Disentanglement with Tri-stage Training for Noise-Robust Speech Recognition
por: Chen, Shuangyuan, et al.
Publicado: (2025)
por: Chen, Shuangyuan, et al.
Publicado: (2025)
Trade-offs Between Capacity and Robustness in Neural Audio Codecs for Adversarially Robust Speech Recognition
por: Prescott, Jordan, et al.
Publicado: (2026)
por: Prescott, Jordan, et al.
Publicado: (2026)
DiCoW: Diarization-Conditioned Whisper for Target Speaker Automatic Speech Recognition
por: Polok, Alexander, et al.
Publicado: (2024)
por: Polok, Alexander, et al.
Publicado: (2024)
Improving Whispered Speech Recognition Performance using Pseudo-whispered based Data Augmentation
por: Lin, Zhaofeng, et al.
Publicado: (2023)
por: Lin, Zhaofeng, et al.
Publicado: (2023)
Robust LLM-based Audio-Visual Speech Recognition with Sparse Modality Alignment and Visual Unit-Guided Refinement
por: Su, Fei, et al.
Publicado: (2026)
por: Su, Fei, et al.
Publicado: (2026)
DCIM-AVSR : Efficient Audio-Visual Speech Recognition via Dual Conformer Interaction Module
por: Wang, Xinyu, et al.
Publicado: (2024)
por: Wang, Xinyu, et al.
Publicado: (2024)
XLAVS-R: Cross-Lingual Audio-Visual Speech Representation Learning for Noise-Robust Speech Perception
por: Han, HyoJung, et al.
Publicado: (2024)
por: Han, HyoJung, et al.
Publicado: (2024)
Prompting Whisper for Joint Speech Transcription and Diarization
por: Zamyrova, Mariia, et al.
Publicado: (2026)
por: Zamyrova, Mariia, et al.
Publicado: (2026)
Probing Whisper for Dysarthric Speech in Detection and Assessment
por: Yue, Zhengjun, et al.
Publicado: (2025)
por: Yue, Zhengjun, et al.
Publicado: (2025)
A Transcription Prompt-based Efficient Audio Large Language Model for Robust Speech Recognition
por: Li, Yangze, et al.
Publicado: (2024)
por: Li, Yangze, et al.
Publicado: (2024)
Classification of Spontaneous and Scripted Speech for Multilingual Audio
por: Elisha, Shahar, et al.
Publicado: (2024)
por: Elisha, Shahar, et al.
Publicado: (2024)
USAD: Universal Speech and Audio Representation via Distillation
por: Chang, Heng-Jui, et al.
Publicado: (2025)
por: Chang, Heng-Jui, et al.
Publicado: (2025)
Investigation of Whisper ASR Hallucinations Induced by Non-Speech Audio
por: Barański, Mateusz, et al.
Publicado: (2025)
por: Barański, Mateusz, et al.
Publicado: (2025)
LCB-net: Long-Context Biasing for Audio-Visual Speech Recognition
por: Yu, Fan, et al.
Publicado: (2024)
por: Yu, Fan, et al.
Publicado: (2024)
A Neural Speech Codec for Noise Robust Speech Coding
por: Huang, Jiayi, et al.
Publicado: (2023)
por: Huang, Jiayi, et al.
Publicado: (2023)
Robust Audio-Visual Speech Enhancement: Correcting Misassignments in Complex Environments with Advanced Post-Processing
por: Ren, Wenze, et al.
Publicado: (2024)
por: Ren, Wenze, et al.
Publicado: (2024)
Audio Flamingo Next: Next-Generation Open Audio-Language Models for Speech, Sound, and Music
por: Ghosh, Sreyan, et al.
Publicado: (2026)
por: Ghosh, Sreyan, et al.
Publicado: (2026)
Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities
por: Kong, Zhifeng, et al.
Publicado: (2024)
por: Kong, Zhifeng, et al.
Publicado: (2024)
WhispEar: A Bi-directional Framework for Scaling Whispered Speech Conversion via Pseudo-Parallel Whisper Generation
por: Fang, Zihao, et al.
Publicado: (2026)
por: Fang, Zihao, et al.
Publicado: (2026)
Swedish Whispers; Leveraging a Massive Speech Corpus for Swedish Speech Recognition
por: Vesterbacka, Leonora, et al.
Publicado: (2025)
por: Vesterbacka, Leonora, et al.
Publicado: (2025)
Robust Speech Recognition with Schrödinger Bridge-Based Speech Enhancement
por: Nasretdinov, Rauf, et al.
Publicado: (2025)
por: Nasretdinov, Rauf, et al.
Publicado: (2025)
The TEA-ASLP System for Multilingual Conversational Speech Recognition and Speech Diarization in MLC-SLM 2025 Challenge
por: Xue, Hongfei, et al.
Publicado: (2025)
por: Xue, Hongfei, et al.
Publicado: (2025)
Multilingual DistilWhisper: Efficient Distillation of Multi-task Speech Models via Language-Specific Experts
por: Ferraz, Thomas Palmeira, et al.
Publicado: (2023)
por: Ferraz, Thomas Palmeira, et al.
Publicado: (2023)
Improving Rare-Word Recognition of Whisper in Zero-Shot Settings
por: Jogi, Yash, et al.
Publicado: (2025)
por: Jogi, Yash, et al.
Publicado: (2025)
Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition
por: Wang, Shih-heng, et al.
Publicado: (2024)
por: Wang, Shih-heng, et al.
Publicado: (2024)
Ejemplares similares
-
Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation
por: Rouditchenko, Andrew, et al.
Publicado: (2024) -
Omni-R1: Do You Really Need Audio to Fine-Tune Your Audio LLM?
por: Rouditchenko, Andrew, et al.
Publicado: (2025) -
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment
por: Araujo, Edson, et al.
Publicado: (2025) -
Towards Audio Token Compression in Large Audio Language Models
por: Bhati, Saurabhchand, et al.
Publicado: (2025) -
DASS: Distilled Audio State Space Models Are Stronger and More Duration-Scalable Learners
por: Bhati, Saurabhchand, et al.
Publicado: (2024)