Cocktail-Party Audio-Visual Speech Recognition
Fuente:
arXiv
Guardado en:
| Autores principales: | Nguyen, Thai-Binh, Pham, Ngoc-Quan, Waibel, Alexander |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
A Cocktail-Party Benchmark: Multi-Modal dataset and Comparative Evaluation Results
por: Nguyen, Thai-Binh, et al.
Publicado: (2025)
por: Nguyen, Thai-Binh, et al.
Publicado: (2025)
Weight Factorization and Centralization for Continual Learning in Speech Recognition
por: Ugan, Enes Yavuz, et al.
Publicado: (2025)
por: Ugan, Enes Yavuz, et al.
Publicado: (2025)
MSA-ASR: Efficient Multilingual Speaker Attribution with frozen ASR Models
por: Nguyen, Thai-Binh, et al.
Publicado: (2024)
por: Nguyen, Thai-Binh, et al.
Publicado: (2024)
Convoifilter: A case study of doing cocktail party speech recognition
por: Nguyen, Thai-Binh, et al.
Publicado: (2023)
por: Nguyen, Thai-Binh, et al.
Publicado: (2023)
Beyond Transcripts: A Renewed Perspective on Audio Chaptering
por: Retkowski, Fabian, et al.
Publicado: (2026)
por: Retkowski, Fabian, et al.
Publicado: (2026)
Accent conversion using discrete units with parallel data synthesized from controllable accented TTS
por: Nguyen, Tuan Nam, et al.
Publicado: (2024)
por: Nguyen, Tuan Nam, et al.
Publicado: (2024)
Adapting Language Balance in Code-Switching Speech
por: Ugan, Enes Yavuz, et al.
Publicado: (2025)
por: Ugan, Enes Yavuz, et al.
Publicado: (2025)
Towards continually learning new languages
por: Pham, Ngoc-Quan, et al.
Publicado: (2022)
por: Pham, Ngoc-Quan, et al.
Publicado: (2022)
Streaming Non-Autoregressive Model for Accent Conversion and Pronunciation Improvement
por: Nguyen, Tuan-Nam, et al.
Publicado: (2025)
por: Nguyen, Tuan-Nam, et al.
Publicado: (2025)
Bayesian Low-Rank Factorization for Robust Model Adaptation
por: Ugan, Enes Yavuz, et al.
Publicado: (2025)
por: Ugan, Enes Yavuz, et al.
Publicado: (2025)
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition
por: Nguyen, Thai-Binh, et al.
Publicado: (2025)
por: Nguyen, Thai-Binh, et al.
Publicado: (2025)
Lombard Speech Synthesis for Any Voice with Controllable Style Embeddings
por: Akti, Seymanur, et al.
Publicado: (2026)
por: Akti, Seymanur, et al.
Publicado: (2026)
Direct Preference Optimization for English-Mandarin Code-Switching Speech Recognition in Audio LLMs
por: Quang, Trung Nguyen, et al.
Publicado: (2026)
por: Quang, Trung Nguyen, et al.
Publicado: (2026)
SARA: Stress Test Reasoning in Audio Deepfake Detection
por: Nguyen, Binh, et al.
Publicado: (2026)
por: Nguyen, Binh, et al.
Publicado: (2026)
Blending LLMs into Cascaded Speech Translation: KIT's Offline Speech Translation System for IWSLT 2024
por: Koneru, Sai, et al.
Publicado: (2024)
por: Koneru, Sai, et al.
Publicado: (2024)
Multi-Channel Speech Enhancement for Cocktail Party Speech Emotion Recognition
por: Chen, Youjun, et al.
Publicado: (2026)
por: Chen, Youjun, et al.
Publicado: (2026)
Predictive Speech Recognition and End-of-Utterance Detection Towards Spoken Dialog Systems
por: Zink, Oswald, et al.
Publicado: (2024)
por: Zink, Oswald, et al.
Publicado: (2024)
What You Read Isn't What You Hear: Linguistic Sensitivity in Deepfake Speech Detection
por: Nguyen, Binh, et al.
Publicado: (2025)
por: Nguyen, Binh, et al.
Publicado: (2025)
Conversational Speech Recognition by Learning Audio-textual Cross-modal Contextual Representation
por: Wei, Kun, et al.
Publicado: (2023)
por: Wei, Kun, et al.
Publicado: (2023)
Summarizing Speech: A Comprehensive Survey
por: Retkowski, Fabian, et al.
Publicado: (2025)
por: Retkowski, Fabian, et al.
Publicado: (2025)
Improving Pronunciation and Accent Conversion through Knowledge Distillation And Synthetic Ground-Truth from Native TTS
por: Nguyen, Tuan Nam, et al.
Publicado: (2024)
por: Nguyen, Tuan Nam, et al.
Publicado: (2024)
DialogueAgents: A Hybrid Agent-Based Speech Synthesis Framework for Multi-Party Dialogue
por: Li, Xiang, et al.
Publicado: (2025)
por: Li, Xiang, et al.
Publicado: (2025)
Decoupled Vocabulary Learning Enables Zero-Shot Translation from Unseen Languages
por: Mullov, Carlos, et al.
Publicado: (2024)
por: Mullov, Carlos, et al.
Publicado: (2024)
MultiMed: Multilingual Medical Speech Recognition via Attention Encoder Decoder
por: Le-Duc, Khai, et al.
Publicado: (2024)
por: Le-Duc, Khai, et al.
Publicado: (2024)
XLAVS-R: Cross-Lingual Audio-Visual Speech Representation Learning for Noise-Robust Speech Perception
por: Han, HyoJung, et al.
Publicado: (2024)
por: Han, HyoJung, et al.
Publicado: (2024)
Zipper-LoRA: Dynamic Parameter Decoupling for Speech-LLM based Multilingual Speech Recognition
por: Mei, Yuxiang, et al.
Publicado: (2026)
por: Mei, Yuxiang, et al.
Publicado: (2026)
Improving Audio-Visual Speech Recognition by Lip-Subword Correlation Based Visual Pre-training and Cross-Modal Fusion Encoder
por: Dai, Yusheng, et al.
Publicado: (2023)
por: Dai, Yusheng, et al.
Publicado: (2023)
USAD: Universal Speech and Audio Representation via Distillation
por: Chang, Heng-Jui, et al.
Publicado: (2025)
por: Chang, Heng-Jui, et al.
Publicado: (2025)
LinTO Audio and Textual Datasets to Train and Evaluate Automatic Speech Recognition in Tunisian Arabic Dialect
por: Naouara, Hedi, et al.
Publicado: (2025)
por: Naouara, Hedi, et al.
Publicado: (2025)
DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech with Factorized Discrete Flow Matching
por: Nguyen, Ngoc-Son, et al.
Publicado: (2025)
por: Nguyen, Ngoc-Son, et al.
Publicado: (2025)
Explainable Transformer-CNN Fusion for Noise-Robust Speech Emotion Recognition
por: Chakrabarty, Sudip, et al.
Publicado: (2025)
por: Chakrabarty, Sudip, et al.
Publicado: (2025)
Towards Orthographically-Informed Evaluation of Speech Recognition Systems for Indian Languages
por: Bhogale, Kaushal Santosh, et al.
Publicado: (2026)
por: Bhogale, Kaushal Santosh, et al.
Publicado: (2026)
UniMoE-Audio: Unified Speech and Music Generation with Dynamic-Capacity MoE
por: Liu, Zhenyu, et al.
Publicado: (2025)
por: Liu, Zhenyu, et al.
Publicado: (2025)
ELEGANCE: Efficient LLM Guidance for Audio-Visual Target Speech Extraction
por: Wu, Wenxuan, et al.
Publicado: (2025)
por: Wu, Wenxuan, et al.
Publicado: (2025)
Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation
por: Rahimi, Akam, et al.
Publicado: (2025)
por: Rahimi, Akam, et al.
Publicado: (2025)
KIT's Offline Speech Translation and Instruction Following Submission for IWSLT 2025
por: Koneru, Sai, et al.
Publicado: (2025)
por: Koneru, Sai, et al.
Publicado: (2025)
MSR-86K: An Evolving, Multilingual Corpus with 86,300 Hours of Transcribed Audio for Speech Recognition Research
por: Li, Song, et al.
Publicado: (2024)
por: Li, Song, et al.
Publicado: (2024)
Learning More with Less: Self-Supervised Approaches for Low-Resource Speech Emotion Recognition
por: Gong, Ziwei, et al.
Publicado: (2025)
por: Gong, Ziwei, et al.
Publicado: (2025)
Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio
por: He, Xinlu, et al.
Publicado: (2025)
por: He, Xinlu, et al.
Publicado: (2025)
Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation
por: Goncalves, Lucas, et al.
Publicado: (2024)
por: Goncalves, Lucas, et al.
Publicado: (2024)
Ejemplares similares
-
A Cocktail-Party Benchmark: Multi-Modal dataset and Comparative Evaluation Results
por: Nguyen, Thai-Binh, et al.
Publicado: (2025) -
Weight Factorization and Centralization for Continual Learning in Speech Recognition
por: Ugan, Enes Yavuz, et al.
Publicado: (2025) -
MSA-ASR: Efficient Multilingual Speaker Attribution with frozen ASR Models
por: Nguyen, Thai-Binh, et al.
Publicado: (2024) -
Convoifilter: A case study of doing cocktail party speech recognition
por: Nguyen, Thai-Binh, et al.
Publicado: (2023) -
Beyond Transcripts: A Renewed Perspective on Audio Chaptering
por: Retkowski, Fabian, et al.
Publicado: (2026)