Enhancing CTC-Based Visual Speech Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Laux, Hendrik, Schmeink, Anke |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multilingual Audio-Visual Speech Recognition with Hybrid CTC/RNN-T Fast Conformer
von: Burchi, Maxime, et al.
Veröffentlicht: (2024)
von: Burchi, Maxime, et al.
Veröffentlicht: (2024)
Efficient Training for Multilingual Visual Speech Recognition: Pre-training with Discretized Visual Speech Representation
von: Kim, Minsu, et al.
Veröffentlicht: (2024)
von: Kim, Minsu, et al.
Veröffentlicht: (2024)
Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation
von: Rouditchenko, Andrew, et al.
Veröffentlicht: (2024)
von: Rouditchenko, Andrew, et al.
Veröffentlicht: (2024)
MoME: Mixture of Matryoshka Experts for Audio-Visual Speech Recognition
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2025)
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2025)
CNVSRC 2024: The Second Chinese Continuous Visual Speech Recognition Challenge
von: Liu, Zehua, et al.
Veröffentlicht: (2025)
von: Liu, Zehua, et al.
Veröffentlicht: (2025)
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition
von: Rouditchenko, Andrew, et al.
Veröffentlicht: (2025)
von: Rouditchenko, Andrew, et al.
Veröffentlicht: (2025)
Mitigating Attention Sinks and Massive Activations in Audio-Visual Speech Recognition with LLMs
von: Anand, et al.
Veröffentlicht: (2025)
von: Anand, et al.
Veröffentlicht: (2025)
Adaptive Audio-Visual Speech Recognition via Matryoshka-Based Multimodal LLMs
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2025)
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2025)
VAPO: End-to-end Slide-Enhanced Speech Recognition with Omni-modal Large Language Models
von: Hu, Rui, et al.
Veröffentlicht: (2025)
von: Hu, Rui, et al.
Veröffentlicht: (2025)
AlignVSR: Audio-Visual Cross-Modal Alignment for Visual Speech Recognition
von: Liu, Zehua, et al.
Veröffentlicht: (2024)
von: Liu, Zehua, et al.
Veröffentlicht: (2024)
Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing
von: Liu, Zehua, et al.
Veröffentlicht: (2025)
von: Liu, Zehua, et al.
Veröffentlicht: (2025)
Dr. SHAP-AV: Decoding Relative Modality Contributions via Shapley Attribution in Audio-Visual Speech Recognition
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2026)
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2026)
Large Language Models are Strong Audio-Visual Speech Recognition Learners
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2024)
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2024)
UniCUE: Unified Recognition and Generation Framework for Chinese Cued Speech Video-to-Speech Generation
von: Wang, Jinting, et al.
Veröffentlicht: (2025)
von: Wang, Jinting, et al.
Veröffentlicht: (2025)
Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech Representations
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2025)
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2025)
MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition with Minimal Multimodal Speech Tokens
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2025)
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2025)
SlideAVSR: A Dataset of Paper Explanation Videos for Audio-Visual Speech Recognition
von: Wang, Hao, et al.
Veröffentlicht: (2024)
von: Wang, Hao, et al.
Veröffentlicht: (2024)
AD-AVSR: Asymmetric Dual-stream Enhancement for Robust Audio-Visual Speech Recognition
von: Xue, Junxiao, et al.
Veröffentlicht: (2025)
von: Xue, Junxiao, et al.
Veröffentlicht: (2025)
Computation and Parameter Efficient Multi-Modal Fusion Transformer for Cued Speech Recognition
von: Liu, Lei, et al.
Veröffentlicht: (2024)
von: Liu, Lei, et al.
Veröffentlicht: (2024)
Omni-AVSR: Towards Unified Multimodal Speech Recognition with Large Language Models
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2025)
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2025)
Seeing Speech and Sound: Distinguishing and Locating Audios in Visual Scenes
von: Ryu, Hyeonggon, et al.
Veröffentlicht: (2025)
von: Ryu, Hyeonggon, et al.
Veröffentlicht: (2025)
SwinLip: An Efficient Visual Speech Encoder for Lip Reading Using Swin Transformer
von: Park, Young-Hu, et al.
Veröffentlicht: (2025)
von: Park, Young-Hu, et al.
Veröffentlicht: (2025)
RTFS-Net: Recurrent Time-Frequency Modelling for Efficient Audio-Visual Speech Separation
von: Pegg, Samuel, et al.
Veröffentlicht: (2023)
von: Pegg, Samuel, et al.
Veröffentlicht: (2023)
Emotional Vietnamese Speech-Based Depression Diagnosis Using Dynamic Attention Mechanism
von: D., Quang-Anh N., et al.
Veröffentlicht: (2024)
von: D., Quang-Anh N., et al.
Veröffentlicht: (2024)
AISHELL6-whisper: A Chinese Mandarin Audio-visual Whisper Speech Dataset with Speech Recognition Baselines
von: Li, Cancan, et al.
Veröffentlicht: (2025)
von: Li, Cancan, et al.
Veröffentlicht: (2025)
Dynamic Derivation and Elimination: Audio Visual Segmentation with Enhanced Audio Semantics
von: Liu, Chen, et al.
Veröffentlicht: (2025)
von: Liu, Chen, et al.
Veröffentlicht: (2025)
Emotional Face-to-Speech
von: Ye, Jiaxin, et al.
Veröffentlicht: (2025)
von: Ye, Jiaxin, et al.
Veröffentlicht: (2025)
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2024)
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2024)
United we stand, Divided we fall: Handling Weak Complementary Relationships for Audio-Visual Emotion Recognition in Valence-Arousal Space
von: Praveen, R. Gnana, et al.
Veröffentlicht: (2025)
von: Praveen, R. Gnana, et al.
Veröffentlicht: (2025)
Robust Audiovisual Speech Recognition Models with Mixture-of-Experts
von: Wu, Yihan, et al.
Veröffentlicht: (2024)
von: Wu, Yihan, et al.
Veröffentlicht: (2024)
Towards Reliable Audio Deepfake Attribution and Model Recognition: A Multi-Level Autoencoder-Based Framework
von: Di Pierno, Andrea, et al.
Veröffentlicht: (2025)
von: Di Pierno, Andrea, et al.
Veröffentlicht: (2025)
Unimodal Aggregation for CTC-based Speech Recognition
von: Fang, Ying, et al.
Veröffentlicht: (2023)
von: Fang, Ying, et al.
Veröffentlicht: (2023)
From Detection to Correction: Backdoor-Resilient Face Recognition via Vision-Language Trigger Detection and Noise-Based Neutralization
von: Wahida, Farah, et al.
Veröffentlicht: (2025)
von: Wahida, Farah, et al.
Veröffentlicht: (2025)
Global-Local Distillation Network-Based Audio-Visual Speaker Tracking with Incomplete Modalities
von: Li, Yidi, et al.
Veröffentlicht: (2024)
von: Li, Yidi, et al.
Veröffentlicht: (2024)
CTC Blank Triggered Dynamic Layer-Skipping for Efficient CTC-based Speech Recognition
von: Hou, Junfeng, et al.
Veröffentlicht: (2024)
von: Hou, Junfeng, et al.
Veröffentlicht: (2024)
IIANet: An Intra- and Inter-Modality Attention Network for Audio-Visual Speech Separation
von: Li, Kai, et al.
Veröffentlicht: (2023)
von: Li, Kai, et al.
Veröffentlicht: (2023)
Audio-Visual Speech Enhancement In Complex Scenarios With Separation And Dereverberation Joint Modeling
von: Du, Jiarong, et al.
Veröffentlicht: (2025)
von: Du, Jiarong, et al.
Veröffentlicht: (2025)
Input Conditioned Layer Dropping in Speech Foundation Models
von: Hannan, Abdul, et al.
Veröffentlicht: (2025)
von: Hannan, Abdul, et al.
Veröffentlicht: (2025)
Spiking Structured State Space Model for Monaural Speech Enhancement
von: Du, Yu, et al.
Veröffentlicht: (2023)
von: Du, Yu, et al.
Veröffentlicht: (2023)
Speaker-Distinguishable CTC: Learning Speaker Distinction Using CTC for Multi-Talker Speech Recognition
von: Sakuma, Asahi, et al.
Veröffentlicht: (2025)
von: Sakuma, Asahi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Multilingual Audio-Visual Speech Recognition with Hybrid CTC/RNN-T Fast Conformer
von: Burchi, Maxime, et al.
Veröffentlicht: (2024) -
Efficient Training for Multilingual Visual Speech Recognition: Pre-training with Discretized Visual Speech Representation
von: Kim, Minsu, et al.
Veröffentlicht: (2024) -
Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation
von: Rouditchenko, Andrew, et al.
Veröffentlicht: (2024) -
MoME: Mixture of Matryoshka Experts for Audio-Visual Speech Recognition
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2025) -
CNVSRC 2024: The Second Chinese Continuous Visual Speech Recognition Challenge
von: Liu, Zehua, et al.
Veröffentlicht: (2025)