Robust LLM-based Audio-Visual Speech Recognition with Sparse Modality Alignment and Visual Unit-Guided Refinement
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Su, Fei, Li, Cancan, Liu, Juan, Ju, Wei, Suo, Hongbin, Li, Ming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AISHELL6-whisper: A Chinese Mandarin Audio-visual Whisper Speech Dataset with Speech Recognition Baselines
von: Li, Cancan, et al.
Veröffentlicht: (2025)
von: Li, Cancan, et al.
Veröffentlicht: (2025)
Human-Inspired Computing for Robust and Efficient Audio-Visual Speech Recognition
von: Liu, Qianhui, et al.
Veröffentlicht: (2024)
von: Liu, Qianhui, et al.
Veröffentlicht: (2024)
AlignVSR: Audio-Visual Cross-Modal Alignment for Visual Speech Recognition
von: Liu, Zehua, et al.
Veröffentlicht: (2024)
von: Liu, Zehua, et al.
Veröffentlicht: (2024)
Robust Wake Word Spotting With Frame-Level Cross-Modal Attention Based Audio-Visual Conformer
von: Wang, Haoxu, et al.
Veröffentlicht: (2024)
von: Wang, Haoxu, et al.
Veröffentlicht: (2024)
Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities
von: Sudarsanam, Parthasaarathy, et al.
Veröffentlicht: (2025)
von: Sudarsanam, Parthasaarathy, et al.
Veröffentlicht: (2025)
LCB-net: Long-Context Biasing for Audio-Visual Speech Recognition
von: Yu, Fan, et al.
Veröffentlicht: (2024)
von: Yu, Fan, et al.
Veröffentlicht: (2024)
RAVSS: Robust Audio-Visual Speech Separation in Multi-Speaker Scenarios with Missing Visual Cues
von: Pan, Tianrui, et al.
Veröffentlicht: (2024)
von: Pan, Tianrui, et al.
Veröffentlicht: (2024)
Efficient Video to Audio Mapper with Visual Scene Detection
von: Yi, Mingjing, et al.
Veröffentlicht: (2024)
von: Yi, Mingjing, et al.
Veröffentlicht: (2024)
Rhythmic Foley: A Framework For Seamless Audio-Visual Alignment In Video-to-Audio Synthesis
von: Huang, Zhiqi, et al.
Veröffentlicht: (2024)
von: Huang, Zhiqi, et al.
Veröffentlicht: (2024)
Audio-Visual Speech Separation via Bottleneck Iterative Network
von: Zhang, Sidong, et al.
Veröffentlicht: (2025)
von: Zhang, Sidong, et al.
Veröffentlicht: (2025)
Listening and Seeing Again: Generative Error Correction for Audio-Visual Speech Recognition
von: Liu, Rui, et al.
Veröffentlicht: (2025)
von: Liu, Rui, et al.
Veröffentlicht: (2025)
ELEGANCE: Efficient LLM Guidance for Audio-Visual Target Speech Extraction
von: Wu, Wenxuan, et al.
Veröffentlicht: (2025)
von: Wu, Wenxuan, et al.
Veröffentlicht: (2025)
Bridging The Multi-Modality Gaps of Audio, Visual and Linguistic for Speech Enhancement
von: Lin, Meng-Ping, et al.
Veröffentlicht: (2025)
von: Lin, Meng-Ping, et al.
Veröffentlicht: (2025)
Zero-Shot Fake Video Detection by Audio-Visual Consistency
von: Li, Xiaolou, et al.
Veröffentlicht: (2024)
von: Li, Xiaolou, et al.
Veröffentlicht: (2024)
AD-AVSR: Asymmetric Dual-stream Enhancement for Robust Audio-Visual Speech Recognition
von: Xue, Junxiao, et al.
Veröffentlicht: (2025)
von: Xue, Junxiao, et al.
Veröffentlicht: (2025)
AVE Speech: A Comprehensive Multi-Modal Dataset for Speech Recognition Integrating Audio, Visual, and Electromyographic Signals
von: Zhou, Dongliang, et al.
Veröffentlicht: (2025)
von: Zhou, Dongliang, et al.
Veröffentlicht: (2025)
Multi-Task Corrupted Prediction for Learning Robust Audio-Visual Speech Representation
von: Kim, Sungnyun, et al.
Veröffentlicht: (2025)
von: Kim, Sungnyun, et al.
Veröffentlicht: (2025)
Mel-Refine: A Plug-and-Play Approach to Refine Mel-Spectrogram in Audio Generation
von: Guo, Hongming, et al.
Veröffentlicht: (2024)
von: Guo, Hongming, et al.
Veröffentlicht: (2024)
Building Audio-Visual Digital Twins with Smartphones
von: Lan, Zitong, et al.
Veröffentlicht: (2025)
von: Lan, Zitong, et al.
Veröffentlicht: (2025)
Cinematic Audio Source Separation Using Visual Cues
von: Zhang, Kang, et al.
Veröffentlicht: (2026)
von: Zhang, Kang, et al.
Veröffentlicht: (2026)
IIANet: An Intra- and Inter-Modality Attention Network for Audio-Visual Speech Separation
von: Li, Kai, et al.
Veröffentlicht: (2023)
von: Li, Kai, et al.
Veröffentlicht: (2023)
Purification Before Fusion: Toward Mask-Free Speech Enhancement for Robust Audio-Visual Speech Recognition
von: Wu, Linzhi, et al.
Veröffentlicht: (2026)
von: Wu, Linzhi, et al.
Veröffentlicht: (2026)
Audio-Visual Speaker Tracking: Progress, Challenges, and Future Directions
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2023)
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2023)
Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions
von: Yuan, Yi, et al.
Veröffentlicht: (2024)
von: Yuan, Yi, et al.
Veröffentlicht: (2024)
Audio-Visual Speech Enhancement In Complex Scenarios With Separation And Dereverberation Joint Modeling
von: Du, Jiarong, et al.
Veröffentlicht: (2025)
von: Du, Jiarong, et al.
Veröffentlicht: (2025)
CatchPhrase: EXPrompt-Guided Encoder Adaptation for Audio-to-Image Generation
von: Oh, Hyunwoo, et al.
Veröffentlicht: (2025)
von: Oh, Hyunwoo, et al.
Veröffentlicht: (2025)
Quality-Aware End-to-End Audio-Visual Neural Speaker Diarization
von: He, Mao-Kui, et al.
Veröffentlicht: (2024)
von: He, Mao-Kui, et al.
Veröffentlicht: (2024)
MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition with Minimal Multimodal Speech Tokens
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2025)
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2025)
SONIQUE: Video Background Music Generation Using Unpaired Audio-Visual Data
von: Zhang, Liqian, et al.
Veröffentlicht: (2024)
von: Zhang, Liqian, et al.
Veröffentlicht: (2024)
Plug-and-Steer: Decoupling Separation and Selection in Audio-Visual Target Speaker Extraction
von: Kwak, Doyeop, et al.
Veröffentlicht: (2026)
von: Kwak, Doyeop, et al.
Veröffentlicht: (2026)
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction
von: Wu, Wenxuan, et al.
Veröffentlicht: (2025)
von: Wu, Wenxuan, et al.
Veröffentlicht: (2025)
MA-AVT: Modality Alignment for Parameter-Efficient Audio-Visual Transformers
von: Mahmud, Tanvir, et al.
Veröffentlicht: (2024)
von: Mahmud, Tanvir, et al.
Veröffentlicht: (2024)
STA-V2A: Video-to-Audio Generation with Semantic and Temporal Alignment
von: Ren, Yong, et al.
Veröffentlicht: (2024)
von: Ren, Yong, et al.
Veröffentlicht: (2024)
Large Language Models are Strong Audio-Visual Speech Recognition Learners
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2024)
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2024)
A Study of Dropout-Induced Modality Bias on Robustness to Missing Video Frames for Audio-Visual Speech Recognition
von: Dai, Yusheng, et al.
Veröffentlicht: (2024)
von: Dai, Yusheng, et al.
Veröffentlicht: (2024)
RVCBench: Benchmarking the Robustness of Voice Cloning Across Modern Audio Generation Models
von: Jin, Ruinan, et al.
Veröffentlicht: (2026)
von: Jin, Ruinan, et al.
Veröffentlicht: (2026)
Adaptive Audio-Visual Speech Recognition via Matryoshka-Based Multimodal LLMs
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2025)
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2025)
Quantitative Analysis of Audio-Visual Tasks: An Information-Theoretic Perspective
von: Chen, Chen, et al.
Veröffentlicht: (2024)
von: Chen, Chen, et al.
Veröffentlicht: (2024)
VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport
von: Rho, Kyeongha, et al.
Veröffentlicht: (2025)
von: Rho, Kyeongha, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
AISHELL6-whisper: A Chinese Mandarin Audio-visual Whisper Speech Dataset with Speech Recognition Baselines
von: Li, Cancan, et al.
Veröffentlicht: (2025) -
Human-Inspired Computing for Robust and Efficient Audio-Visual Speech Recognition
von: Liu, Qianhui, et al.
Veröffentlicht: (2024) -
AlignVSR: Audio-Visual Cross-Modal Alignment for Visual Speech Recognition
von: Liu, Zehua, et al.
Veröffentlicht: (2024) -
Robust Wake Word Spotting With Frame-Level Cross-Modal Attention Based Audio-Visual Conformer
von: Wang, Haoxu, et al.
Veröffentlicht: (2024) -
Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities
von: Sudarsanam, Parthasaarathy, et al.
Veröffentlicht: (2025)