Adaptive Speaker Embedding Self-Augmentation for Personal Voice Activity Detection with Short Enrollment Speech
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Feng, Fuyuan, Zhang, Wenbin, Gao, Yu, Xu, Longting, Mou, Xiaofeng, Xu, Yi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multi-modal Speech Enhancement with Limited Electromyography Channels
von: Feng, Fuyuan, et al.
Veröffentlicht: (2025)
von: Feng, Fuyuan, et al.
Veröffentlicht: (2025)
EEND-SAA: Enrollment-Less Main Speaker Voice Activity Detection Using Self-Attention Attractors
von: Wu, Wen-Yung, et al.
Veröffentlicht: (2025)
von: Wu, Wen-Yung, et al.
Veröffentlicht: (2025)
Unified Architecture and Unsupervised Speech Disentanglement for Speaker Embedding-Free Enrollment in Personalized Speech Enhancement
von: Huang, Ziling, et al.
Veröffentlicht: (2025)
von: Huang, Ziling, et al.
Veröffentlicht: (2025)
Universal Speaker Embedding Free Target Speaker Extraction and Personal Voice Activity Detection
von: Zeng, Bang, et al.
Veröffentlicht: (2025)
von: Zeng, Bang, et al.
Veröffentlicht: (2025)
Speaker Embeddings With Weakly Supervised Voice Activity Detection For Efficient Speaker Diarization
von: Thienpondt, Jenthe, et al.
Veröffentlicht: (2024)
von: Thienpondt, Jenthe, et al.
Veröffentlicht: (2024)
Unmixing the Crowd: Learning Mixture-to-Set Speaker Embeddings for Enrollment-Free Target Speech Extraction
von: Sidharth, FNU, et al.
Veröffentlicht: (2026)
von: Sidharth, FNU, et al.
Veröffentlicht: (2026)
NanoVoice: Efficient Speaker-Adaptive Text-to-Speech for Multiple Speakers
von: Park, Nohil, et al.
Veröffentlicht: (2024)
von: Park, Nohil, et al.
Veröffentlicht: (2024)
Retrieval-Augmented Self-Taught Reasoning Model with Adaptive Chain-of-Thought for ASR Named Entity Correction
von: An, Junjie, et al.
Veröffentlicht: (2026)
von: An, Junjie, et al.
Veröffentlicht: (2026)
End-to-End Direction-Aware Keyword Spotting with Spatial Priors in Noisy Environments
von: Wang, Rui, et al.
Veröffentlicht: (2026)
von: Wang, Rui, et al.
Veröffentlicht: (2026)
Device Feature based on Graph Fourier Transformation with Logarithmic Processing For Detection of Replay Speech Attacks
von: He, Mingrui, et al.
Veröffentlicht: (2024)
von: He, Mingrui, et al.
Veröffentlicht: (2024)
EvoTSE: Evolving Enrollment for Target Speaker Extraction
von: Liu, Zikai, et al.
Veröffentlicht: (2026)
von: Liu, Zikai, et al.
Veröffentlicht: (2026)
Robust Speech Activity Detection in the Presence of Singing Voice
von: Grundhuber, Philipp, et al.
Veröffentlicht: (2025)
von: Grundhuber, Philipp, et al.
Veröffentlicht: (2025)
Voice Conversion Augmentation for Speaker Recognition on Defective Datasets
von: Tao, Ruijie, et al.
Veröffentlicht: (2024)
von: Tao, Ruijie, et al.
Veröffentlicht: (2024)
Profile-Error-Tolerant Target-Speaker Voice Activity Detection
von: Wang, Dongmei, et al.
Veröffentlicht: (2023)
von: Wang, Dongmei, et al.
Veröffentlicht: (2023)
Exploring Speech Foundation Models for Speaker Diarization Across Lifespan
von: Xu, Anfeng, et al.
Veröffentlicht: (2026)
von: Xu, Anfeng, et al.
Veröffentlicht: (2026)
Post-Training Embedding Alignment for Decoupling Enrollment and Runtime Speaker Recognition Models
von: Gao, Chenyang, et al.
Veröffentlicht: (2024)
von: Gao, Chenyang, et al.
Veröffentlicht: (2024)
Multi-Input Multi-Output Target-Speaker Voice Activity Detection For Unified, Flexible, and Robust Audio-Visual Speaker Diarization
von: Cheng, Ming, et al.
Veröffentlicht: (2024)
von: Cheng, Ming, et al.
Veröffentlicht: (2024)
Adversarial Attacks and Robust Defenses in Speaker Embedding based Zero-Shot Text-to-Speech System
von: Li, Ze, et al.
Veröffentlicht: (2024)
von: Li, Ze, et al.
Veröffentlicht: (2024)
GLOBE: A High-quality English Corpus with Global Accents for Zero-shot Speaker Adaptive Text-to-Speech
von: Wang, Wenbin, et al.
Veröffentlicht: (2024)
von: Wang, Wenbin, et al.
Veröffentlicht: (2024)
Single-Microphone Speaker Separation and Voice Activity Detection in Noisy and Reverberant Environments
von: Opochinsky, Renana, et al.
Veröffentlicht: (2024)
von: Opochinsky, Renana, et al.
Veröffentlicht: (2024)
Joint Training of Speaker Embedding Extractor, Speech and Overlap Detection for Diarization
von: Pálka, Petr, et al.
Veröffentlicht: (2024)
von: Pálka, Petr, et al.
Veröffentlicht: (2024)
End-to-End Target Speaker Speech Recognition Using Context-Aware Attention Mechanisms for Challenging Enrollment Scenario
von: Ghane, Mohsen, et al.
Veröffentlicht: (2025)
von: Ghane, Mohsen, et al.
Veröffentlicht: (2025)
Emotional Styles Hide in Deep Speaker Embeddings: Disentangle Deep Speaker Embeddings for Speaker Clustering
von: Lin, Chaohao, et al.
Veröffentlicht: (2025)
von: Lin, Chaohao, et al.
Veröffentlicht: (2025)
Residual Speaker Representation for One-Shot Voice Conversion
von: Xu, Le, et al.
Veröffentlicht: (2023)
von: Xu, Le, et al.
Veröffentlicht: (2023)
Adaptive Data Augmentation with NaturalSpeech3 for Far-field Speaker Verification
von: Zhang, Li, et al.
Veröffentlicht: (2025)
von: Zhang, Li, et al.
Veröffentlicht: (2025)
Flow-TSVAD: Target-Speaker Voice Activity Detection via Latent Flow Matching
von: Chen, Zhengyang, et al.
Veröffentlicht: (2024)
von: Chen, Zhengyang, et al.
Veröffentlicht: (2024)
VoiceGuider: Enhancing Out-of-Domain Performance in Parameter-Efficient Speaker-Adaptive Text-to-Speech via Autoguidance
von: Yeom, Jiheum, et al.
Veröffentlicht: (2024)
von: Yeom, Jiheum, et al.
Veröffentlicht: (2024)
A Comprehensive Investigation on Speaker Augmentation for Speaker Recognition
von: Zhou, Zhenyu, et al.
Veröffentlicht: (2024)
von: Zhou, Zhenyu, et al.
Veröffentlicht: (2024)
End-to-End DOA-Guided Speech Extraction in Noisy Multi-Talker Scenarios
von: Jing, Kangqi, et al.
Veröffentlicht: (2025)
von: Jing, Kangqi, et al.
Veröffentlicht: (2025)
What Does the Speaker Embedding Encode?
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
Attacking Voice Anonymization Systems with Augmented Feature and Speaker Identity Difference
von: Zhang, Yanzhe, et al.
Veröffentlicht: (2024)
von: Zhang, Yanzhe, et al.
Veröffentlicht: (2024)
Target Speaker Extraction through Comparing Noisy Positive and Negative Audio Enrollments
von: Xu, Shitong, et al.
Veröffentlicht: (2025)
von: Xu, Shitong, et al.
Veröffentlicht: (2025)
Creating Personalized Synthetic Voices from Articulation Impaired Speech Using Augmented Reconstruction Loss
von: Tian, Yusheng, et al.
Veröffentlicht: (2024)
von: Tian, Yusheng, et al.
Veröffentlicht: (2024)
End-to-End Integration of Speech Emotion Recognition with Voice Activity Detection using Self-Supervised Learning Features
von: Yamashita, Natsuo, et al.
Veröffentlicht: (2024)
von: Yamashita, Natsuo, et al.
Veröffentlicht: (2024)
DualStream Contextual Fusion Network: Efficient Target Speaker Extraction by Leveraging Mixture and Enrollment Interactions
von: Xue, Ke, et al.
Veröffentlicht: (2025)
von: Xue, Ke, et al.
Veröffentlicht: (2025)
AS-Speech: Adaptive Style For Speech Synthesis
von: Li, Zhipeng, et al.
Veröffentlicht: (2024)
von: Li, Zhipeng, et al.
Veröffentlicht: (2024)
Spoken Language Corpora Augmentation with Domain-Specific Voice-Cloned Speech
von: Czyżnikiewicz, Mateusz, et al.
Veröffentlicht: (2024)
von: Czyżnikiewicz, Mateusz, et al.
Veröffentlicht: (2024)
Channel-Combination Algorithms for Robust Distant Voice Activity and Overlapped Speech Detection
von: Mariotte, Théo, et al.
Veröffentlicht: (2024)
von: Mariotte, Théo, et al.
Veröffentlicht: (2024)
Robust Audio-Visual Target Speaker Extraction with Emotion-Aware Multiple Enrollment Fusion
von: Jin, Zhan, et al.
Veröffentlicht: (2025)
von: Jin, Zhan, et al.
Veröffentlicht: (2025)
PAS-SE: Personalized Auxiliary-Sensor Speech Enhancement for Voice Pickup in Hearables
von: Ohlenbusch, Mattes, et al.
Veröffentlicht: (2025)
von: Ohlenbusch, Mattes, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Multi-modal Speech Enhancement with Limited Electromyography Channels
von: Feng, Fuyuan, et al.
Veröffentlicht: (2025) -
EEND-SAA: Enrollment-Less Main Speaker Voice Activity Detection Using Self-Attention Attractors
von: Wu, Wen-Yung, et al.
Veröffentlicht: (2025) -
Unified Architecture and Unsupervised Speech Disentanglement for Speaker Embedding-Free Enrollment in Personalized Speech Enhancement
von: Huang, Ziling, et al.
Veröffentlicht: (2025) -
Universal Speaker Embedding Free Target Speaker Extraction and Personal Voice Activity Detection
von: Zeng, Bang, et al.
Veröffentlicht: (2025) -
Speaker Embeddings With Weakly Supervised Voice Activity Detection For Efficient Speaker Diarization
von: Thienpondt, Jenthe, et al.
Veröffentlicht: (2024)