Efficient Audio-Visual Speech Separation with Discrete Lip Semantics and Multi-Scale Global-Local Attention
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Kai, Gao, Kejun, Hu, Xiaolin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RTFS-Net: Recurrent Time-Frequency Modelling for Efficient Audio-Visual Speech Separation
von: Pegg, Samuel, et al.
Veröffentlicht: (2023)
von: Pegg, Samuel, et al.
Veröffentlicht: (2023)
An Audio-Visual Speech Separation Model Inspired by Cortico-Thalamo-Cortical Circuits
von: Li, Kai, et al.
Veröffentlicht: (2022)
von: Li, Kai, et al.
Veröffentlicht: (2022)
IIANet: An Intra- and Inter-Modality Attention Network for Audio-Visual Speech Separation
von: Li, Kai, et al.
Veröffentlicht: (2023)
von: Li, Kai, et al.
Veröffentlicht: (2023)
SwinLip: An Efficient Visual Speech Encoder for Lip Reading Using Swin Transformer
von: Park, Young-Hu, et al.
Veröffentlicht: (2025)
von: Park, Young-Hu, et al.
Veröffentlicht: (2025)
Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation
von: Goncalves, Lucas, et al.
Veröffentlicht: (2024)
von: Goncalves, Lucas, et al.
Veröffentlicht: (2024)
Semantic Audio-Visual Navigation in Continuous Environments
von: Zeng, Yichen, et al.
Veröffentlicht: (2026)
von: Zeng, Yichen, et al.
Veröffentlicht: (2026)
VAEmo: Efficient Representation Learning for Visual-Audio Emotion with Knowledge Injection
von: Cheng, Hao, et al.
Veröffentlicht: (2025)
von: Cheng, Hao, et al.
Veröffentlicht: (2025)
Audio-Visual Speech Enhancement In Complex Scenarios With Separation And Dereverberation Joint Modeling
von: Du, Jiarong, et al.
Veröffentlicht: (2025)
von: Du, Jiarong, et al.
Veröffentlicht: (2025)
Improving Sound Source Localization with Joint Slot Attention on Image and Audio
von: Kim, Inho, et al.
Veröffentlicht: (2025)
von: Kim, Inho, et al.
Veröffentlicht: (2025)
Robust Audio-Visual Segmentation via Audio-Guided Visual Convergent Alignment
von: Liu, Chen, et al.
Veröffentlicht: (2025)
von: Liu, Chen, et al.
Veröffentlicht: (2025)
Efficient Training for Multilingual Visual Speech Recognition: Pre-training with Discretized Visual Speech Representation
von: Kim, Minsu, et al.
Veröffentlicht: (2024)
von: Kim, Minsu, et al.
Veröffentlicht: (2024)
Mitigating Attention Sinks and Massive Activations in Audio-Visual Speech Recognition with LLMs
von: Anand, et al.
Veröffentlicht: (2025)
von: Anand, et al.
Veröffentlicht: (2025)
LTA-L2S: Lexical Tone-Aware Lip-to-Speech Synthesis for Mandarin with Cross-Lingual Transfer Learning
von: Yang, Kang, et al.
Veröffentlicht: (2025)
von: Yang, Kang, et al.
Veröffentlicht: (2025)
Semantic Noise Reduction via Teacher-Guided Dual-Path Audio-Visual Representation Learning
von: Wang, Linge, et al.
Veröffentlicht: (2026)
von: Wang, Linge, et al.
Veröffentlicht: (2026)
Dynamic Derivation and Elimination: Audio Visual Segmentation with Enhanced Audio Semantics
von: Liu, Chen, et al.
Veröffentlicht: (2025)
von: Liu, Chen, et al.
Veröffentlicht: (2025)
Global-Local Distillation Network-Based Audio-Visual Speaker Tracking with Incomplete Modalities
von: Li, Yidi, et al.
Veröffentlicht: (2024)
von: Li, Yidi, et al.
Veröffentlicht: (2024)
Pilot-guided Multimodal Semantic Communication for Audio-Visual Event Localization
von: Yu, Fei, et al.
Veröffentlicht: (2024)
von: Yu, Fei, et al.
Veröffentlicht: (2024)
Separate in the Speech Chain: Cross-Modal Conditional Audio-Visual Target Speech Extraction
von: Mu, Zhaoxi, et al.
Veröffentlicht: (2024)
von: Mu, Zhaoxi, et al.
Veröffentlicht: (2024)
Towards Accurate Lip-to-Speech Synthesis in-the-Wild
von: Hegde, Sindhu, et al.
Veröffentlicht: (2024)
von: Hegde, Sindhu, et al.
Veröffentlicht: (2024)
Cross-Modal Binary Attention: An Energy-Efficient Fusion Framework for Audio-Visual Learning
von: Saleh, Mohamed, et al.
Veröffentlicht: (2026)
von: Saleh, Mohamed, et al.
Veröffentlicht: (2026)
AlignVSR: Audio-Visual Cross-Modal Alignment for Visual Speech Recognition
von: Liu, Zehua, et al.
Veröffentlicht: (2024)
von: Liu, Zehua, et al.
Veröffentlicht: (2024)
Semantics-Aware Human Motion Generation from Audio Instructions
von: Wang, Zi-An, et al.
Veröffentlicht: (2025)
von: Wang, Zi-An, et al.
Veröffentlicht: (2025)
DDAVS: Disentangled Audio Semantics and Delayed Bidirectional Alignment for Audio-Visual Segmentation
von: Tian, Jingqi, et al.
Veröffentlicht: (2025)
von: Tian, Jingqi, et al.
Veröffentlicht: (2025)
MoLT: Mixture of Layer-Wise Tokens for Efficient Audio-Visual Learning
von: Rho, Kyeongha, et al.
Veröffentlicht: (2025)
von: Rho, Kyeongha, et al.
Veröffentlicht: (2025)
Speech Audio Generation from dynamic MRI via a Knowledge Enhanced Conditional Variational Autoencoder
von: Li, Yaxuan, et al.
Veröffentlicht: (2025)
von: Li, Yaxuan, et al.
Veröffentlicht: (2025)
AV-Edit: Multimodal Generative Sound Effect Editing via Audio-Visual Semantic Joint Control
von: Guo, Xinyue, et al.
Veröffentlicht: (2025)
von: Guo, Xinyue, et al.
Veröffentlicht: (2025)
Segment Beyond View: Handling Partially Missing Modality for Audio-Visual Semantic Segmentation
von: Wu, Renjie, et al.
Veröffentlicht: (2023)
von: Wu, Renjie, et al.
Veröffentlicht: (2023)
Seeing Speech and Sound: Distinguishing and Locating Audios in Visual Scenes
von: Ryu, Hyeonggon, et al.
Veröffentlicht: (2025)
von: Ryu, Hyeonggon, et al.
Veröffentlicht: (2025)
MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition with Minimal Multimodal Speech Tokens
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2025)
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2025)
AV-RIR: Audio-Visual Room Impulse Response Estimation
von: Ratnarajah, Anton, et al.
Veröffentlicht: (2023)
von: Ratnarajah, Anton, et al.
Veröffentlicht: (2023)
MultiSoundGen: Video-to-Audio Generation for Multi-Event Scenarios via SlowFast Contrastive Audio-Visual Pretraining and Direct Preference Optimization
von: Yang, Jianxuan, et al.
Veröffentlicht: (2025)
von: Yang, Jianxuan, et al.
Veröffentlicht: (2025)
Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation
von: Rouditchenko, Andrew, et al.
Veröffentlicht: (2024)
von: Rouditchenko, Andrew, et al.
Veröffentlicht: (2024)
Seeing Soundscapes: Audio-Visual Generation and Separation from Soundscapes Using Audio-Visual Separator
von: Kang, Minjae, et al.
Veröffentlicht: (2025)
von: Kang, Minjae, et al.
Veröffentlicht: (2025)
MoME: Mixture of Matryoshka Experts for Audio-Visual Speech Recognition
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2025)
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2025)
SonoWorld: From One Image to a 3D Audio-Visual Scene
von: Jin, Derong, et al.
Veröffentlicht: (2026)
von: Jin, Derong, et al.
Veröffentlicht: (2026)
Toward Fine-Grained Speech Inpainting Forensics:A Dataset, Method, and Metric for Multi-Region Tampering Localization
von: Vu, Tung, et al.
Veröffentlicht: (2026)
von: Vu, Tung, et al.
Veröffentlicht: (2026)
Rethinking Audio-Visual Adversarial Vulnerability from Temporal and Modality Perspectives
von: Zhang, Zeliang, et al.
Veröffentlicht: (2025)
von: Zhang, Zeliang, et al.
Veröffentlicht: (2025)
WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM
von: Tang, Changli, et al.
Veröffentlicht: (2025)
von: Tang, Changli, et al.
Veröffentlicht: (2025)
DASH: Dynamic Audio-Driven Semantic Chunking for Efficient Omnimodal Token Compression
von: Li, Bingzhou, et al.
Veröffentlicht: (2026)
von: Li, Bingzhou, et al.
Veröffentlicht: (2026)
OLKAVS: An Open Large-Scale Korean Audio-Visual Speech Dataset
von: Park, Jeongkyun, et al.
Veröffentlicht: (2023)
von: Park, Jeongkyun, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
RTFS-Net: Recurrent Time-Frequency Modelling for Efficient Audio-Visual Speech Separation
von: Pegg, Samuel, et al.
Veröffentlicht: (2023) -
An Audio-Visual Speech Separation Model Inspired by Cortico-Thalamo-Cortical Circuits
von: Li, Kai, et al.
Veröffentlicht: (2022) -
IIANet: An Intra- and Inter-Modality Attention Network for Audio-Visual Speech Separation
von: Li, Kai, et al.
Veröffentlicht: (2023) -
SwinLip: An Efficient Visual Speech Encoder for Lip Reading Using Swin Transformer
von: Park, Young-Hu, et al.
Veröffentlicht: (2025) -
Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation
von: Goncalves, Lucas, et al.
Veröffentlicht: (2024)