Two-stage Audio-Visual Target Speaker Extraction System for Real-Time Processing On Edge Device
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Zixuan, Zhang, Xueliang, Miao, Lei, Yan, Zhipeng, Sun, Ying, Zhu, Chong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Robust Target Speaker Direction of Arrival Estimation
von: Li, Zixuan, et al.
Veröffentlicht: (2024)
von: Li, Zixuan, et al.
Veröffentlicht: (2024)
Online Audio-Visual Autoregressive Speaker Extraction
von: Pan, Zexu, et al.
Veröffentlicht: (2025)
von: Pan, Zexu, et al.
Veröffentlicht: (2025)
3S-TSE: Efficient Three-Stage Target Speaker Extraction for Real-Time and Low-Resource Applications
von: He, Shulin, et al.
Veröffentlicht: (2023)
von: He, Shulin, et al.
Veröffentlicht: (2023)
Audio-Visual Target Speaker Extraction with Reverse Selective Auditory Attention
von: Tao, Ruijie, et al.
Veröffentlicht: (2024)
von: Tao, Ruijie, et al.
Veröffentlicht: (2024)
Listen to Extract: Onset-Prompted Target Speaker Extraction
von: Shen, Pengjie, et al.
Veröffentlicht: (2025)
von: Shen, Pengjie, et al.
Veröffentlicht: (2025)
Robust Audio-Visual Target Speaker Extraction with Emotion-Aware Multiple Enrollment Fusion
von: Jin, Zhan, et al.
Veröffentlicht: (2025)
von: Jin, Zhan, et al.
Veröffentlicht: (2025)
Target Speaker Extraction with Curriculum Learning
von: Liu, Yun, et al.
Veröffentlicht: (2024)
von: Liu, Yun, et al.
Veröffentlicht: (2024)
Plug-and-Steer: Decoupling Separation and Selection in Audio-Visual Target Speaker Extraction
von: Kwak, Doyeop, et al.
Veröffentlicht: (2026)
von: Kwak, Doyeop, et al.
Veröffentlicht: (2026)
Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data
von: Liu, Yun, et al.
Veröffentlicht: (2024)
von: Liu, Yun, et al.
Veröffentlicht: (2024)
Multi-Level Speaker Representation for Target Speaker Extraction
von: Zhang, Ke, et al.
Veröffentlicht: (2024)
von: Zhang, Ke, et al.
Veröffentlicht: (2024)
HRTF-guided Binaural Target Speaker Extraction with Real-World Validation
von: Ellinson, Yoav, et al.
Veröffentlicht: (2026)
von: Ellinson, Yoav, et al.
Veröffentlicht: (2026)
Enhancing Target Speaker Extraction with Explicit Speaker Consistency Modeling
von: Wu, Shu, et al.
Veröffentlicht: (2025)
von: Wu, Shu, et al.
Veröffentlicht: (2025)
Plug-and-Play Co-Occurring Face Attention for Robust Audio-Visual Speaker Extraction
von: Pan, Zexu, et al.
Veröffentlicht: (2025)
von: Pan, Zexu, et al.
Veröffentlicht: (2025)
USEF-TSE: Universal Speaker Embedding Free Target Speaker Extraction
von: Zeng, Bang, et al.
Veröffentlicht: (2024)
von: Zeng, Bang, et al.
Veröffentlicht: (2024)
Beyond Speaker Identity: Text Guided Target Speech Extraction
von: Huo, Mingyue, et al.
Veröffentlicht: (2025)
von: Huo, Mingyue, et al.
Veröffentlicht: (2025)
Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR
von: Ma, Hao, et al.
Veröffentlicht: (2025)
von: Ma, Hao, et al.
Veröffentlicht: (2025)
Improved Feature Extraction Network for Neuro-Oriented Target Speaker Extraction
von: Fan, Cunhang, et al.
Veröffentlicht: (2025)
von: Fan, Cunhang, et al.
Veröffentlicht: (2025)
On the effectiveness of enrollment speech augmentation for Target Speaker Extraction
von: Li, Junjie, et al.
Veröffentlicht: (2024)
von: Li, Junjie, et al.
Veröffentlicht: (2024)
Binaural Target Speaker Extraction using Individualized HRTF
von: Ellinson, Yoav, et al.
Veröffentlicht: (2025)
von: Ellinson, Yoav, et al.
Veröffentlicht: (2025)
Universal Speaker Embedding Free Target Speaker Extraction and Personal Voice Activity Detection
von: Zeng, Bang, et al.
Veröffentlicht: (2025)
von: Zeng, Bang, et al.
Veröffentlicht: (2025)
AVFSNet: Audio-Visual Speech Separation for Flexible Number of Speakers with Multi-Scale and Multi-Task Learning
von: Zhang, Daning, et al.
Veröffentlicht: (2025)
von: Zhang, Daning, et al.
Veröffentlicht: (2025)
Target Speaker Extraction through Comparing Noisy Positive and Negative Audio Enrollments
von: Xu, Shitong, et al.
Veröffentlicht: (2025)
von: Xu, Shitong, et al.
Veröffentlicht: (2025)
$C^2$AV-TSE: Context and Confidence-aware Audio Visual Target Speaker Extraction
von: Wu, Wenxuan, et al.
Veröffentlicht: (2025)
von: Wu, Wenxuan, et al.
Veröffentlicht: (2025)
Two-stage Framework for Robust Speech Emotion Recognition Using Target Speaker Extraction in Human Speech Noise Conditions
von: Mi, Jinyi, et al.
Veröffentlicht: (2024)
von: Mi, Jinyi, et al.
Veröffentlicht: (2024)
Utilizing Speaker Profiles for Impersonation Audio Detection
von: Gu, Hao, et al.
Veröffentlicht: (2024)
von: Gu, Hao, et al.
Veröffentlicht: (2024)
Inter-Speaker Relative Cues for Text-Guided Target Speech Extraction
von: Dai, Wang, et al.
Veröffentlicht: (2025)
von: Dai, Wang, et al.
Veröffentlicht: (2025)
Discriminative-Generative Target Speaker Extraction with Decoder-Only Language Models
von: Zeng, Bang, et al.
Veröffentlicht: (2026)
von: Zeng, Bang, et al.
Veröffentlicht: (2026)
DualStream Contextual Fusion Network: Efficient Target Speaker Extraction by Leveraging Mixture and Enrollment Interactions
von: Xue, Ke, et al.
Veröffentlicht: (2025)
von: Xue, Ke, et al.
Veröffentlicht: (2025)
SoloAudio: Target Sound Extraction with Language-oriented Audio Diffusion Transformer
von: Wang, Helin, et al.
Veröffentlicht: (2024)
von: Wang, Helin, et al.
Veröffentlicht: (2024)
Audio-Visual Speaker Tracking: Progress, Challenges, and Future Directions
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2023)
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2023)
pTSE-T: Presentation Target Speaker Extraction using Unaligned Text Cues
von: Jiang, Ziyang, et al.
Veröffentlicht: (2024)
von: Jiang, Ziyang, et al.
Veröffentlicht: (2024)
SQ-Whisper: Speaker-Querying based Whisper Model for Target-Speaker ASR
von: Guo, Pengcheng, et al.
Veröffentlicht: (2024)
von: Guo, Pengcheng, et al.
Veröffentlicht: (2024)
M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction
von: Fan, Cunhang, et al.
Veröffentlicht: (2025)
von: Fan, Cunhang, et al.
Veröffentlicht: (2025)
An Investigation on Speaker Augmentation for End-to-End Speaker Extraction
von: You, Zhenghai, et al.
Veröffentlicht: (2025)
von: You, Zhenghai, et al.
Veröffentlicht: (2025)
Hyperdimensional Intelligent Sensing for Efficient Real-Time Audio Processing on Extreme Edge
von: Yun, Sanggeon, et al.
Veröffentlicht: (2025)
von: Yun, Sanggeon, et al.
Veröffentlicht: (2025)
Audiosockets: A Python socket package for Real-Time Audio Processing
von: Shu, Nicolas, et al.
Veröffentlicht: (2024)
von: Shu, Nicolas, et al.
Veröffentlicht: (2024)
WeSep: A Scalable and Flexible Toolkit Towards Generalizable Target Speaker Extraction
von: Wang, Shuai, et al.
Veröffentlicht: (2024)
von: Wang, Shuai, et al.
Veröffentlicht: (2024)
Review of MEMS Speakers for Audio Applications
von: Wittek, Nils, et al.
Veröffentlicht: (2025)
von: Wittek, Nils, et al.
Veröffentlicht: (2025)
Leveraging Audio-Only Data for Text-Queried Target Sound Extraction
von: Saijo, Kohei, et al.
Veröffentlicht: (2024)
von: Saijo, Kohei, et al.
Veröffentlicht: (2024)
Unmixing the Crowd: Learning Mixture-to-Set Speaker Embeddings for Enrollment-Free Target Speech Extraction
von: Sidharth, FNU, et al.
Veröffentlicht: (2026)
von: Sidharth, FNU, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Robust Target Speaker Direction of Arrival Estimation
von: Li, Zixuan, et al.
Veröffentlicht: (2024) -
Online Audio-Visual Autoregressive Speaker Extraction
von: Pan, Zexu, et al.
Veröffentlicht: (2025) -
3S-TSE: Efficient Three-Stage Target Speaker Extraction for Real-Time and Low-Resource Applications
von: He, Shulin, et al.
Veröffentlicht: (2023) -
Audio-Visual Target Speaker Extraction with Reverse Selective Auditory Attention
von: Tao, Ruijie, et al.
Veröffentlicht: (2024) -
Listen to Extract: Onset-Prompted Target Speaker Extraction
von: Shen, Pengjie, et al.
Veröffentlicht: (2025)