Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Zehua, Li, Xiaolou, Guo, Li, Li, Lantian, Wang, Dong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CNVSRC 2024: The Second Chinese Continuous Visual Speech Recognition Challenge
von: Liu, Zehua, et al.
Veröffentlicht: (2025)
von: Liu, Zehua, et al.
Veröffentlicht: (2025)
AlignVSR: Audio-Visual Cross-Modal Alignment for Visual Speech Recognition
von: Liu, Zehua, et al.
Veröffentlicht: (2024)
von: Liu, Zehua, et al.
Veröffentlicht: (2024)
Zero-Shot Fake Video Detection by Audio-Visual Consistency
von: Li, Xiaolou, et al.
Veröffentlicht: (2024)
von: Li, Xiaolou, et al.
Veröffentlicht: (2024)
Quantitative Analysis of Audio-Visual Tasks: An Information-Theoretic Perspective
von: Chen, Chen, et al.
Veröffentlicht: (2024)
von: Chen, Chen, et al.
Veröffentlicht: (2024)
Large Language Models are Strong Audio-Visual Speech Recognition Learners
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2024)
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2024)
SE/BN Adapter: Parametric Efficient Domain Adaptation for Speaker Recognition
von: Wang, Tianhao, et al.
Veröffentlicht: (2024)
von: Wang, Tianhao, et al.
Veröffentlicht: (2024)
Training-Free Deepfake Voice Recognition by Leveraging Large-Scale Pre-Trained Models
von: Pianese, Alessandro, et al.
Veröffentlicht: (2024)
von: Pianese, Alessandro, et al.
Veröffentlicht: (2024)
BrainWhisperer: Leveraging Large-Scale ASR Models for Neural Speech Decoding
von: Boccato, Tommaso, et al.
Veröffentlicht: (2026)
von: Boccato, Tommaso, et al.
Veröffentlicht: (2026)
Omni-AVSR: Towards Unified Multimodal Speech Recognition with Large Language Models
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2025)
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2025)
UniCUE: Unified Recognition and Generation Framework for Chinese Cued Speech Video-to-Speech Generation
von: Wang, Jinting, et al.
Veröffentlicht: (2025)
von: Wang, Jinting, et al.
Veröffentlicht: (2025)
Streaming Speech Recognition with Decoder-Only Large Language Models and Latency Optimization
von: Wan, Genshun, et al.
Veröffentlicht: (2026)
von: Wan, Genshun, et al.
Veröffentlicht: (2026)
VAPO: End-to-end Slide-Enhanced Speech Recognition with Omni-modal Large Language Models
von: Hu, Rui, et al.
Veröffentlicht: (2025)
von: Hu, Rui, et al.
Veröffentlicht: (2025)
Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
von: Liao, Shijia, et al.
Veröffentlicht: (2024)
von: Liao, Shijia, et al.
Veröffentlicht: (2024)
Dr. SHAP-AV: Decoding Relative Modality Contributions via Shapley Attribution in Audio-Visual Speech Recognition
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2026)
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2026)
RTFS-Net: Recurrent Time-Frequency Modelling for Efficient Audio-Visual Speech Separation
von: Pegg, Samuel, et al.
Veröffentlicht: (2023)
von: Pegg, Samuel, et al.
Veröffentlicht: (2023)
Seeing the Context: Rich Visual Context-Aware Speech Recognition via Multimodal Reasoning
von: Tian, Wenjie, et al.
Veröffentlicht: (2026)
von: Tian, Wenjie, et al.
Veröffentlicht: (2026)
Enhancing CTC-Based Visual Speech Recognition
von: Laux, Hendrik, et al.
Veröffentlicht: (2024)
von: Laux, Hendrik, et al.
Veröffentlicht: (2024)
A Comprehensive Investigation on Speaker Augmentation for Speaker Recognition
von: Zhou, Zhenyu, et al.
Veröffentlicht: (2024)
von: Zhou, Zhenyu, et al.
Veröffentlicht: (2024)
Efficient Training for Multilingual Visual Speech Recognition: Pre-training with Discretized Visual Speech Representation
von: Kim, Minsu, et al.
Veröffentlicht: (2024)
von: Kim, Minsu, et al.
Veröffentlicht: (2024)
Computation and Parameter Efficient Multi-Modal Fusion Transformer for Cued Speech Recognition
von: Liu, Lei, et al.
Veröffentlicht: (2024)
von: Liu, Lei, et al.
Veröffentlicht: (2024)
EmoQ: Speech Emotion Recognition via Speech-Aware Q-Former and Large Language Model
von: Yang, Yiqing, et al.
Veröffentlicht: (2025)
von: Yang, Yiqing, et al.
Veröffentlicht: (2025)
CoLM-DSR: Leveraging Neural Codec Language Modeling for Multi-Modal Dysarthric Speech Reconstruction
von: Chen, Xueyuan, et al.
Veröffentlicht: (2024)
von: Chen, Xueyuan, et al.
Veröffentlicht: (2024)
Oceanship: A Large-Scale Dataset for Underwater Audio Target Recognition
von: Li, Zeyu, et al.
Veröffentlicht: (2024)
von: Li, Zeyu, et al.
Veröffentlicht: (2024)
Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation
von: Rouditchenko, Andrew, et al.
Veröffentlicht: (2024)
von: Rouditchenko, Andrew, et al.
Veröffentlicht: (2024)
USM-Lite: Quantization and Sparsity Aware Fine-tuning for Speech Recognition with Universal Speech Models
von: Ding, Shaojin, et al.
Veröffentlicht: (2023)
von: Ding, Shaojin, et al.
Veröffentlicht: (2023)
MoME: Mixture of Matryoshka Experts for Audio-Visual Speech Recognition
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2025)
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2025)
Audio-Visual Speech Enhancement In Complex Scenarios With Separation And Dereverberation Joint Modeling
von: Du, Jiarong, et al.
Veröffentlicht: (2025)
von: Du, Jiarong, et al.
Veröffentlicht: (2025)
Mamba-based Decoder-Only Approach with Bidirectional Speech Modeling for Speech Recognition
von: Masuyama, Yoshiki, et al.
Veröffentlicht: (2024)
von: Masuyama, Yoshiki, et al.
Veröffentlicht: (2024)
A Transcription Prompt-based Efficient Audio Large Language Model for Robust Speech Recognition
von: Li, Yangze, et al.
Veröffentlicht: (2024)
von: Li, Yangze, et al.
Veröffentlicht: (2024)
Augmenting Polish Automatic Speech Recognition System With Synthetic Data
von: Bondaruk, Łukasz, et al.
Veröffentlicht: (2024)
von: Bondaruk, Łukasz, et al.
Veröffentlicht: (2024)
Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition
von: Bai, Ye, et al.
Veröffentlicht: (2024)
von: Bai, Ye, et al.
Veröffentlicht: (2024)
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition
von: Rouditchenko, Andrew, et al.
Veröffentlicht: (2025)
von: Rouditchenko, Andrew, et al.
Veröffentlicht: (2025)
Mitigating Attention Sinks and Massive Activations in Audio-Visual Speech Recognition with LLMs
von: Anand, et al.
Veröffentlicht: (2025)
von: Anand, et al.
Veröffentlicht: (2025)
Neural Scoring: A Refreshed End-to-End Approach for Speaker Recognition in Complex Conditions
von: Lin, Wan, et al.
Veröffentlicht: (2024)
von: Lin, Wan, et al.
Veröffentlicht: (2024)
Leveraging Self-Supervised Models for Automatic Whispered Speech Recognition
von: Farhadipour, Aref, et al.
Veröffentlicht: (2024)
von: Farhadipour, Aref, et al.
Veröffentlicht: (2024)
Enhancing Speech Large Language Models with Prompt-Aware Mixture of Audio Encoders
von: Shan, Weiqiao, et al.
Veröffentlicht: (2025)
von: Shan, Weiqiao, et al.
Veröffentlicht: (2025)
Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech Representations
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2025)
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2025)
Long-Context Speech Synthesis with Context-Aware Memory
von: Li, Zhipeng, et al.
Veröffentlicht: (2025)
von: Li, Zhipeng, et al.
Veröffentlicht: (2025)
Zero-Shot Recognition of Dysarthric Speech Using Commercial Automatic Speech Recognition and Multimodal Large Language Models
von: Alsayegh, Ali, et al.
Veröffentlicht: (2025)
von: Alsayegh, Ali, et al.
Veröffentlicht: (2025)
Speech-Mamba: Long-Context Speech Recognition with Selective State Spaces Models
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2024)
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
CNVSRC 2024: The Second Chinese Continuous Visual Speech Recognition Challenge
von: Liu, Zehua, et al.
Veröffentlicht: (2025) -
AlignVSR: Audio-Visual Cross-Modal Alignment for Visual Speech Recognition
von: Liu, Zehua, et al.
Veröffentlicht: (2024) -
Zero-Shot Fake Video Detection by Audio-Visual Consistency
von: Li, Xiaolou, et al.
Veröffentlicht: (2024) -
Quantitative Analysis of Audio-Visual Tasks: An Information-Theoretic Perspective
von: Chen, Chen, et al.
Veröffentlicht: (2024) -
Large Language Models are Strong Audio-Visual Speech Recognition Learners
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2024)