Plug-and-Steer: Decoupling Separation and Selection in Audio-Visual Target Speaker Extraction
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kwak, Doyeop, Lee, Suyeon, Chung, Joon Son |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LRS-VoxMM: A benchmark for in-the-wild audio-visual speech recognition
von: Kwak, Doyeop, et al.
Veröffentlicht: (2026)
von: Kwak, Doyeop, et al.
Veröffentlicht: (2026)
Cinematic Audio Source Separation Using Visual Cues
von: Zhang, Kang, et al.
Veröffentlicht: (2026)
von: Zhang, Kang, et al.
Veröffentlicht: (2026)
RAVSS: Robust Audio-Visual Speech Separation in Multi-Speaker Scenarios with Missing Visual Cues
von: Pan, Tianrui, et al.
Veröffentlicht: (2024)
von: Pan, Tianrui, et al.
Veröffentlicht: (2024)
Model-Guided Dual-Role Alignment for High-Fidelity Open-Domain Video-to-Audio Generation
von: Zhang, Kang, et al.
Veröffentlicht: (2025)
von: Zhang, Kang, et al.
Veröffentlicht: (2025)
LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport
von: Rho, Kyeongha, et al.
Veröffentlicht: (2025)
von: Rho, Kyeongha, et al.
Veröffentlicht: (2025)
Quality-Aware End-to-End Audio-Visual Neural Speaker Diarization
von: He, Mao-Kui, et al.
Veröffentlicht: (2024)
von: He, Mao-Kui, et al.
Veröffentlicht: (2024)
Audio-Visual Speaker Tracking: Progress, Challenges, and Future Directions
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2023)
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2023)
$C^2$AV-TSE: Context and Confidence-aware Audio Visual Target Speaker Extraction
von: Wu, Wenxuan, et al.
Veröffentlicht: (2025)
von: Wu, Wenxuan, et al.
Veröffentlicht: (2025)
SpeakerLLM: A Speaker-Specialized Audio-LLM for Speaker Understanding and Verification Reasoning
von: Nam, KiHyun, et al.
Veröffentlicht: (2026)
von: Nam, KiHyun, et al.
Veröffentlicht: (2026)
pTSE-T: Presentation Target Speaker Extraction using Unaligned Text Cues
von: Jiang, Ziyang, et al.
Veröffentlicht: (2024)
von: Jiang, Ziyang, et al.
Veröffentlicht: (2024)
Audio-Visual Speech Separation via Bottleneck Iterative Network
von: Zhang, Sidong, et al.
Veröffentlicht: (2025)
von: Zhang, Sidong, et al.
Veröffentlicht: (2025)
Mel-Refine: A Plug-and-Play Approach to Refine Mel-Spectrogram in Audio Generation
von: Guo, Hongming, et al.
Veröffentlicht: (2024)
von: Guo, Hongming, et al.
Veröffentlicht: (2024)
ELEGANCE: Efficient LLM Guidance for Audio-Visual Target Speech Extraction
von: Wu, Wenxuan, et al.
Veröffentlicht: (2025)
von: Wu, Wenxuan, et al.
Veröffentlicht: (2025)
FlowAVSE: Efficient Audio-Visual Speech Enhancement with Conditional Flow Matching
von: Jung, Chaeyoung, et al.
Veröffentlicht: (2024)
von: Jung, Chaeyoung, et al.
Veröffentlicht: (2024)
EDNet: A Versatile Speech Enhancement Framework with Gating Mamba Mechanism and Phase Shift-Invariant Training
von: Kwak, Doyeop, et al.
Veröffentlicht: (2025)
von: Kwak, Doyeop, et al.
Veröffentlicht: (2025)
Spectron: Target Speaker Extraction using Conditional Transformer with Adversarial Refinement
von: Bandyopadhyay, Tathagata
Veröffentlicht: (2024)
von: Bandyopadhyay, Tathagata
Veröffentlicht: (2024)
Audio-Visual Speaker Diarization: Current Databases, Approaches and Challenges
von: Mingote, Victoria, et al.
Veröffentlicht: (2024)
von: Mingote, Victoria, et al.
Veröffentlicht: (2024)
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction
von: Wu, Wenxuan, et al.
Veröffentlicht: (2025)
von: Wu, Wenxuan, et al.
Veröffentlicht: (2025)
Rhythmic Foley: A Framework For Seamless Audio-Visual Alignment In Video-to-Audio Synthesis
von: Huang, Zhiqi, et al.
Veröffentlicht: (2024)
von: Huang, Zhiqi, et al.
Veröffentlicht: (2024)
Building Audio-Visual Digital Twins with Smartphones
von: Lan, Zitong, et al.
Veröffentlicht: (2025)
von: Lan, Zitong, et al.
Veröffentlicht: (2025)
ecVoice: Audio Text Extraction and Optimization of Video Based on Idioms Similarity Replacement
von: Lin, Jinwei
Veröffentlicht: (2024)
von: Lin, Jinwei
Veröffentlicht: (2024)
Efficient Video to Audio Mapper with Visual Scene Detection
von: Yi, Mingjing, et al.
Veröffentlicht: (2024)
von: Yi, Mingjing, et al.
Veröffentlicht: (2024)
Zero-Shot Fake Video Detection by Audio-Visual Consistency
von: Li, Xiaolou, et al.
Veröffentlicht: (2024)
von: Li, Xiaolou, et al.
Veröffentlicht: (2024)
Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions
von: Yuan, Yi, et al.
Veröffentlicht: (2024)
von: Yuan, Yi, et al.
Veröffentlicht: (2024)
Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities
von: Sudarsanam, Parthasaarathy, et al.
Veröffentlicht: (2025)
von: Sudarsanam, Parthasaarathy, et al.
Veröffentlicht: (2025)
LCB-net: Long-Context Biasing for Audio-Visual Speech Recognition
von: Yu, Fan, et al.
Veröffentlicht: (2024)
von: Yu, Fan, et al.
Veröffentlicht: (2024)
Human-Inspired Computing for Robust and Efficient Audio-Visual Speech Recognition
von: Liu, Qianhui, et al.
Veröffentlicht: (2024)
von: Liu, Qianhui, et al.
Veröffentlicht: (2024)
Robust LLM-based Audio-Visual Speech Recognition with Sparse Modality Alignment and Visual Unit-Guided Refinement
von: Su, Fei, et al.
Veröffentlicht: (2026)
von: Su, Fei, et al.
Veröffentlicht: (2026)
Aligning Sight and Sound: Advanced Sound Source Localization Through Audio-Visual Alignment
von: Senocak, Arda, et al.
Veröffentlicht: (2024)
von: Senocak, Arda, et al.
Veröffentlicht: (2024)
SONIQUE: Video Background Music Generation Using Unpaired Audio-Visual Data
von: Zhang, Liqian, et al.
Veröffentlicht: (2024)
von: Zhang, Liqian, et al.
Veröffentlicht: (2024)
Robust Wake Word Spotting With Frame-Level Cross-Modal Attention Based Audio-Visual Conformer
von: Wang, Haoxu, et al.
Veröffentlicht: (2024)
von: Wang, Haoxu, et al.
Veröffentlicht: (2024)
Noise-Conditioned Mixture-of-Experts Framework for Robust Speaker Verification
von: Gu, Bin, et al.
Veröffentlicht: (2025)
von: Gu, Bin, et al.
Veröffentlicht: (2025)
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey
von: Su, Yi, et al.
Veröffentlicht: (2025)
von: Su, Yi, et al.
Veröffentlicht: (2025)
Audio-Visual Target Speaker Extraction with Reverse Selective Auditory Attention
von: Tao, Ruijie, et al.
Veröffentlicht: (2024)
von: Tao, Ruijie, et al.
Veröffentlicht: (2024)
CatchPhrase: EXPrompt-Guided Encoder Adaptation for Audio-to-Image Generation
von: Oh, Hyunwoo, et al.
Veröffentlicht: (2025)
von: Oh, Hyunwoo, et al.
Veröffentlicht: (2025)
Face-Voice Association for Audiovisual Active Speaker Detection in Egocentric Recordings
von: Clarke, Jason, et al.
Veröffentlicht: (2025)
von: Clarke, Jason, et al.
Veröffentlicht: (2025)
UNMIXX: Untangling Highly Correlated Singing Voices Mixtures
von: Jung, Jihoo, et al.
Veröffentlicht: (2026)
von: Jung, Jihoo, et al.
Veröffentlicht: (2026)
REWIND: Speech Time Reversal for Enhancing Speaker Representations in Diffusion-based Voice Conversion
von: Biyani, Ishan D., et al.
Veröffentlicht: (2025)
von: Biyani, Ishan D., et al.
Veröffentlicht: (2025)
Preserving Speaker Information in Direct Speech-to-Speech Translation with Non-Autoregressive Generation and Pretraining
von: Zhou, Rui, et al.
Veröffentlicht: (2024)
von: Zhou, Rui, et al.
Veröffentlicht: (2024)
MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues
von: Li, Junjie, et al.
Veröffentlicht: (2024)
von: Li, Junjie, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
LRS-VoxMM: A benchmark for in-the-wild audio-visual speech recognition
von: Kwak, Doyeop, et al.
Veröffentlicht: (2026) -
Cinematic Audio Source Separation Using Visual Cues
von: Zhang, Kang, et al.
Veröffentlicht: (2026) -
RAVSS: Robust Audio-Visual Speech Separation in Multi-Speaker Scenarios with Missing Visual Cues
von: Pan, Tianrui, et al.
Veröffentlicht: (2024) -
Model-Guided Dual-Role Alignment for High-Fidelity Open-Domain Video-to-Audio Generation
von: Zhang, Kang, et al.
Veröffentlicht: (2025) -
LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport
von: Rho, Kyeongha, et al.
Veröffentlicht: (2025)