Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Wenxuan, Wang, Shuai, Wu, Xixin, Meng, Helen, Li, Haizhou |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ELEGANCE: Efficient LLM Guidance for Audio-Visual Target Speech Extraction
von: Wu, Wenxuan, et al.
Veröffentlicht: (2025)
von: Wu, Wenxuan, et al.
Veröffentlicht: (2025)
$C^2$AV-TSE: Context and Confidence-aware Audio Visual Target Speaker Extraction
von: Wu, Wenxuan, et al.
Veröffentlicht: (2025)
von: Wu, Wenxuan, et al.
Veröffentlicht: (2025)
Bridging The Multi-Modality Gaps of Audio, Visual and Linguistic for Speech Enhancement
von: Lin, Meng-Ping, et al.
Veröffentlicht: (2025)
von: Lin, Meng-Ping, et al.
Veröffentlicht: (2025)
Target Speech Extraction with Pre-trained AV-HuBERT and Mask-And-Recover Strategy
von: Wu, Wenxuan, et al.
Veröffentlicht: (2024)
von: Wu, Wenxuan, et al.
Veröffentlicht: (2024)
Speech Separation with Pretrained Frontend to Minimize Domain Mismatch
von: Wang, Wupeng, et al.
Veröffentlicht: (2024)
von: Wang, Wupeng, et al.
Veröffentlicht: (2024)
IML-Spikeformer: Input-aware Multi-Level Spiking Transformer for Speech Processing
von: Song, Zeyang, et al.
Veröffentlicht: (2025)
von: Song, Zeyang, et al.
Veröffentlicht: (2025)
Human-Inspired Computing for Robust and Efficient Audio-Visual Speech Recognition
von: Liu, Qianhui, et al.
Veröffentlicht: (2024)
von: Liu, Qianhui, et al.
Veröffentlicht: (2024)
Multi-Task Corrupted Prediction for Learning Robust Audio-Visual Speech Representation
von: Kim, Sungnyun, et al.
Veröffentlicht: (2025)
von: Kim, Sungnyun, et al.
Veröffentlicht: (2025)
Separate in the Speech Chain: Cross-Modal Conditional Audio-Visual Target Speech Extraction
von: Mu, Zhaoxi, et al.
Veröffentlicht: (2024)
von: Mu, Zhaoxi, et al.
Veröffentlicht: (2024)
Audiopedia: Audio QA with Knowledge
von: Penamakuri, Abhirama Subramanyam, et al.
Veröffentlicht: (2024)
von: Penamakuri, Abhirama Subramanyam, et al.
Veröffentlicht: (2024)
Plug-and-Steer: Decoupling Separation and Selection in Audio-Visual Target Speaker Extraction
von: Kwak, Doyeop, et al.
Veröffentlicht: (2026)
von: Kwak, Doyeop, et al.
Veröffentlicht: (2026)
LSTMSE-Net: Long Short Term Speech Enhancement Network for Audio-visual Speech Enhancement
von: Jain, Arnav, et al.
Veröffentlicht: (2024)
von: Jain, Arnav, et al.
Veröffentlicht: (2024)
RAVSS: Robust Audio-Visual Speech Separation in Multi-Speaker Scenarios with Missing Visual Cues
von: Pan, Tianrui, et al.
Veröffentlicht: (2024)
von: Pan, Tianrui, et al.
Veröffentlicht: (2024)
Purification Before Fusion: Toward Mask-Free Speech Enhancement for Robust Audio-Visual Speech Recognition
von: Wu, Linzhi, et al.
Veröffentlicht: (2026)
von: Wu, Linzhi, et al.
Veröffentlicht: (2026)
Network Bending of Diffusion Models for Audio-Visual Generation
von: Dzwonczyk, Luke, et al.
Veröffentlicht: (2024)
von: Dzwonczyk, Luke, et al.
Veröffentlicht: (2024)
Listening and Seeing Again: Generative Error Correction for Audio-Visual Speech Recognition
von: Liu, Rui, et al.
Veröffentlicht: (2025)
von: Liu, Rui, et al.
Veröffentlicht: (2025)
Spectron: Target Speaker Extraction using Conditional Transformer with Adversarial Refinement
von: Bandyopadhyay, Tathagata
Veröffentlicht: (2024)
von: Bandyopadhyay, Tathagata
Veröffentlicht: (2024)
Cinematic Audio Source Separation Using Visual Cues
von: Zhang, Kang, et al.
Veröffentlicht: (2026)
von: Zhang, Kang, et al.
Veröffentlicht: (2026)
Audio-Visual Speech Separation via Bottleneck Iterative Network
von: Zhang, Sidong, et al.
Veröffentlicht: (2025)
von: Zhang, Sidong, et al.
Veröffentlicht: (2025)
LCB-net: Long-Context Biasing for Audio-Visual Speech Recognition
von: Yu, Fan, et al.
Veröffentlicht: (2024)
von: Yu, Fan, et al.
Veröffentlicht: (2024)
Audio-Visual Speaker Tracking: Progress, Challenges, and Future Directions
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2023)
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2023)
Rhythmic Foley: A Framework For Seamless Audio-Visual Alignment In Video-to-Audio Synthesis
von: Huang, Zhiqi, et al.
Veröffentlicht: (2024)
von: Huang, Zhiqi, et al.
Veröffentlicht: (2024)
AWARE: Audio Watermarking with Adversarial Resistance to Edits
von: Pavlović, Kosta, et al.
Veröffentlicht: (2025)
von: Pavlović, Kosta, et al.
Veröffentlicht: (2025)
Robust LLM-based Audio-Visual Speech Recognition with Sparse Modality Alignment and Visual Unit-Guided Refinement
von: Su, Fei, et al.
Veröffentlicht: (2026)
von: Su, Fei, et al.
Veröffentlicht: (2026)
Just Label the Repeats for In-The-Wild Audio-to-Score Alignment
von: Bukey, Irmak, et al.
Veröffentlicht: (2024)
von: Bukey, Irmak, et al.
Veröffentlicht: (2024)
Leveraging Pre-Trained Models for Multimodal Class-Incremental Learning under Adaptive Fusion
von: Chen, Yukun, et al.
Veröffentlicht: (2025)
von: Chen, Yukun, et al.
Veröffentlicht: (2025)
Multimodal Speech Enhancement Using Burst Propagation
von: Raza, Mohsin, et al.
Veröffentlicht: (2022)
von: Raza, Mohsin, et al.
Veröffentlicht: (2022)
VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
ConsistencyTTA: Accelerating Diffusion-Based Text-to-Audio Generation with Consistency Distillation
von: Bai, Yatong, et al.
Veröffentlicht: (2023)
von: Bai, Yatong, et al.
Veröffentlicht: (2023)
Language Model Based Text-to-Audio Generation: Anti-Causally Aligned Collaborative Residual Transformers
von: Wang, Juncheng, et al.
Veröffentlicht: (2025)
von: Wang, Juncheng, et al.
Veröffentlicht: (2025)
Efficient Feature Extraction and Late Fusion Strategy for Audiovisual Emotional Mimicry Intensity Estimation
von: Yu, Jun, et al.
Veröffentlicht: (2024)
von: Yu, Jun, et al.
Veröffentlicht: (2024)
Multimodal Emotion Coupling via Speech-to-Facial and Bodily Gestures in Dyadic Interaction
von: Herbuela, Von Ralph Dane Marquez, et al.
Veröffentlicht: (2025)
von: Herbuela, Von Ralph Dane Marquez, et al.
Veröffentlicht: (2025)
A Simple but Strong Baseline for Sounding Video Generation: Effective Adaptation of Audio and Video Diffusion Models for Joint Generation
von: Ishii, Masato, et al.
Veröffentlicht: (2024)
von: Ishii, Masato, et al.
Veröffentlicht: (2024)
CoAVT: A Cognition-Inspired Unified Audio-Visual-Text Pre-Training Model for Multimodal Processing
von: Yue, Xianghu, et al.
Veröffentlicht: (2024)
von: Yue, Xianghu, et al.
Veröffentlicht: (2024)
Source Separation of Multi-source Raw Music using a Residual Quantized Variational Autoencoder
von: Berti, Leonardo
Veröffentlicht: (2024)
von: Berti, Leonardo
Veröffentlicht: (2024)
Building Audio-Visual Digital Twins with Smartphones
von: Lan, Zitong, et al.
Veröffentlicht: (2025)
von: Lan, Zitong, et al.
Veröffentlicht: (2025)
Beyond Video-to-SFX: Video to Audio Synthesis with Environmentally Aware Speech
von: Niu, Xinlei, et al.
Veröffentlicht: (2025)
von: Niu, Xinlei, et al.
Veröffentlicht: (2025)
Efficient Speech Watermarking for Speech Synthesis via Progressive Knowledge Distillation
von: Cui, Yang, et al.
Veröffentlicht: (2025)
von: Cui, Yang, et al.
Veröffentlicht: (2025)
ecVoice: Audio Text Extraction and Optimization of Video Based on Idioms Similarity Replacement
von: Lin, Jinwei
Veröffentlicht: (2024)
von: Lin, Jinwei
Veröffentlicht: (2024)
ASK: Adaptive Self-improving Knowledge Framework for Audio Text Retrieval
von: Fu, Siyuan, et al.
Veröffentlicht: (2025)
von: Fu, Siyuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
ELEGANCE: Efficient LLM Guidance for Audio-Visual Target Speech Extraction
von: Wu, Wenxuan, et al.
Veröffentlicht: (2025) -
$C^2$AV-TSE: Context and Confidence-aware Audio Visual Target Speaker Extraction
von: Wu, Wenxuan, et al.
Veröffentlicht: (2025) -
Bridging The Multi-Modality Gaps of Audio, Visual and Linguistic for Speech Enhancement
von: Lin, Meng-Ping, et al.
Veröffentlicht: (2025) -
Target Speech Extraction with Pre-trained AV-HuBERT and Mask-And-Recover Strategy
von: Wu, Wenxuan, et al.
Veröffentlicht: (2024) -
Speech Separation with Pretrained Frontend to Minimize Domain Mismatch
von: Wang, Wupeng, et al.
Veröffentlicht: (2024)