openFEAT: Improving Speaker Identification by Open-set Few-shot Embedding Adaptation with Transformer
Fuente:
arXiv
Guardado en:
| Autores principales: | C, Kishan K, Tan, Zhenning, Chen, Long, Jin, Minho, Han, Eunjung, Stolcke, Andreas, Lee, Chul |
|---|---|
| Formato: | Preprint |
| Publicado: |
2022
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SpeakerRPL v2: Robust Open-set Speaker Identification through Enhanced Few-shot Foundation Tuning and Model Fusion
por: Chen, Zhiyong, et al.
Publicado: (2026)
por: Chen, Zhiyong, et al.
Publicado: (2026)
Adversarial Reweighting for Speaker Verification Fairness
por: Jin, Minho, et al.
Publicado: (2022)
por: Jin, Minho, et al.
Publicado: (2022)
Improving fairness in speaker verification via Group-adapted Fusion Network
por: Shen, Hua, et al.
Publicado: (2022)
por: Shen, Hua, et al.
Publicado: (2022)
Post-Training Embedding Alignment for Decoupling Enrollment and Runtime Speaker Recognition Models
por: Gao, Chenyang, et al.
Publicado: (2024)
por: Gao, Chenyang, et al.
Publicado: (2024)
Improving speaker verification robustness with synthetic emotional utterances
por: Koditala, Nikhil Kumar, et al.
Publicado: (2024)
por: Koditala, Nikhil Kumar, et al.
Publicado: (2024)
Fully Few-shot Class-incremental Audio Classification Using Multi-level Embedding Extractor and Ridge Regression Classifier
por: Si, Yongjie, et al.
Publicado: (2025)
por: Si, Yongjie, et al.
Publicado: (2025)
Target Speaker Lipreading by Audio-Visual Self-Distillation Pretraining and Speaker Adaptation
por: Zhang, Jing-Xuan, et al.
Publicado: (2025)
por: Zhang, Jing-Xuan, et al.
Publicado: (2025)
VoxBlink2: A 100K+ Speaker Recognition Corpus and the Open-Set Speaker-Identification Benchmark
por: Lin, Yuke, et al.
Publicado: (2024)
por: Lin, Yuke, et al.
Publicado: (2024)
Speaker Embeddings to Improve Tracking of Intermittent and Moving Speakers
por: Iatariene, Taous, et al.
Publicado: (2025)
por: Iatariene, Taous, et al.
Publicado: (2025)
Rhythm Features for Speaker Identification
por: Mehlman, Nick, et al.
Publicado: (2025)
por: Mehlman, Nick, et al.
Publicado: (2025)
Enhancing Open-Set Speaker Identification through Rapid Tuning with Speaker Reciprocal Points and Negative Sample
por: Chen, Zhiyong, et al.
Publicado: (2024)
por: Chen, Zhiyong, et al.
Publicado: (2024)
Emotional Styles Hide in Deep Speaker Embeddings: Disentangle Deep Speaker Embeddings for Speaker Clustering
por: Lin, Chaohao, et al.
Publicado: (2025)
por: Lin, Chaohao, et al.
Publicado: (2025)
An Investigation of Reprogramming for Cross-Language Adaptation in Speaker Verification Systems
por: Li, Jingyu, et al.
Publicado: (2024)
por: Li, Jingyu, et al.
Publicado: (2024)
What Does the Speaker Embedding Encode?
por: Wang, Shuai, et al.
Publicado: (2025)
por: Wang, Shuai, et al.
Publicado: (2025)
Whisper Speaker Identification: Leveraging Pre-Trained Multilingual Transformers for Robust Speaker Embeddings
por: Emon, Jakaria Islam, et al.
Publicado: (2025)
por: Emon, Jakaria Islam, et al.
Publicado: (2025)
Guided Speaker Embedding
por: Horiguchi, Shota, et al.
Publicado: (2024)
por: Horiguchi, Shota, et al.
Publicado: (2024)
On-the-fly Routing for Zero-shot MoE Speaker Adaptation of Speech Foundation Models for Dysarthric Speech Recognition
por: HU, Shujie, et al.
Publicado: (2025)
por: HU, Shujie, et al.
Publicado: (2025)
Pretraining Multi-Speaker Identification for Neural Speaker Diarization
por: Horiguchi, Shota, et al.
Publicado: (2025)
por: Horiguchi, Shota, et al.
Publicado: (2025)
Mitigating Non-Target Speaker Bias in Guided Speaker Embedding
por: Horiguchi, Shota, et al.
Publicado: (2025)
por: Horiguchi, Shota, et al.
Publicado: (2025)
PseudoVC: Improving One-shot Voice Conversion with Pseudo Paired Data
por: Cao, Songjun, et al.
Publicado: (2025)
por: Cao, Songjun, et al.
Publicado: (2025)
Speaker-Smoothed kNN Speaker Adaptation for End-to-End ASR
por: Li, Shaojun, et al.
Publicado: (2024)
por: Li, Shaojun, et al.
Publicado: (2024)
Speaker Targeting via Self-Speaker Adaptation for Multi-talker ASR
por: Wang, Weiqing, et al.
Publicado: (2025)
por: Wang, Weiqing, et al.
Publicado: (2025)
UNet-Based Fusion and Exponential Moving Average Adaptation for Noise-Robust Speaker Recognition
por: Gan, Chong-Xin, et al.
Publicado: (2026)
por: Gan, Chong-Xin, et al.
Publicado: (2026)
USEF-TSE: Universal Speaker Embedding Free Target Speaker Extraction
por: Zeng, Bang, et al.
Publicado: (2024)
por: Zeng, Bang, et al.
Publicado: (2024)
Unified Architecture and Unsupervised Speech Disentanglement for Speaker Embedding-Free Enrollment in Personalized Speech Enhancement
por: Huang, Ziling, et al.
Publicado: (2025)
por: Huang, Ziling, et al.
Publicado: (2025)
Interpreting the Dimensions of Speaker Embedding Space
por: Huckvale, Mark
Publicado: (2025)
por: Huckvale, Mark
Publicado: (2025)
Test-Time Adaptation for Speech Enhancement via Domain Invariant Embedding Transformation
por: Raichle, Tobias, et al.
Publicado: (2025)
por: Raichle, Tobias, et al.
Publicado: (2025)
Recursive Attentive Pooling for Extracting Speaker Embeddings from Multi-Speaker Recordings
por: Horiguchi, Shota, et al.
Publicado: (2024)
por: Horiguchi, Shota, et al.
Publicado: (2024)
Speaker Embeddings With Weakly Supervised Voice Activity Detection For Efficient Speaker Diarization
por: Thienpondt, Jenthe, et al.
Publicado: (2024)
por: Thienpondt, Jenthe, et al.
Publicado: (2024)
Fully Few-shot Class-incremental Audio Classification Using Expandable Dual-embedding Extractor
por: Si, Yongjie, et al.
Publicado: (2024)
por: Si, Yongjie, et al.
Publicado: (2024)
A Toolkit for Joint Speaker Diarization and Identification with Application to Speaker-Attributed ASR
por: Morrone, Giovanni, et al.
Publicado: (2024)
por: Morrone, Giovanni, et al.
Publicado: (2024)
SEED: Speaker Embedding Enhancement Diffusion Model
por: Nam, KiHyun, et al.
Publicado: (2025)
por: Nam, KiHyun, et al.
Publicado: (2025)
ASRRL-TTS: Agile Speaker Representation Reinforcement Learning for Text-to-Speech Speaker Adaptation
por: Fu, Ruibo, et al.
Publicado: (2024)
por: Fu, Ruibo, et al.
Publicado: (2024)
Universal Speaker Embedding Free Target Speaker Extraction and Personal Voice Activity Detection
por: Zeng, Bang, et al.
Publicado: (2025)
por: Zeng, Bang, et al.
Publicado: (2025)
The Reasonable Effectiveness of Speaker Embeddings for Violence Detection
por: Jain, Sarthak, et al.
Publicado: (2024)
por: Jain, Sarthak, et al.
Publicado: (2024)
Explaining Speaker and Spoof Embeddings via Probing
por: Liu, Xuechen, et al.
Publicado: (2024)
por: Liu, Xuechen, et al.
Publicado: (2024)
Xi+: Uncertainty Supervision for Robust Speaker Embedding
por: Li, Junjie, et al.
Publicado: (2025)
por: Li, Junjie, et al.
Publicado: (2025)
Cochleagram-based Noise Adapted Speaker Identification System for Distorted Speech
por: Ahmed, Sabbir, et al.
Publicado: (2025)
por: Ahmed, Sabbir, et al.
Publicado: (2025)
DNN based HRIRs Identification with a Continuously Rotating Speaker Array
por: Ko, Byeong-Yun, et al.
Publicado: (2025)
por: Ko, Byeong-Yun, et al.
Publicado: (2025)
Unsupervised Single-Channel Speech Separation with a Diffusion Prior under Speaker-Embedding Guidance
por: Shi, Runwu, et al.
Publicado: (2025)
por: Shi, Runwu, et al.
Publicado: (2025)
Ejemplares similares
-
SpeakerRPL v2: Robust Open-set Speaker Identification through Enhanced Few-shot Foundation Tuning and Model Fusion
por: Chen, Zhiyong, et al.
Publicado: (2026) -
Adversarial Reweighting for Speaker Verification Fairness
por: Jin, Minho, et al.
Publicado: (2022) -
Improving fairness in speaker verification via Group-adapted Fusion Network
por: Shen, Hua, et al.
Publicado: (2022) -
Post-Training Embedding Alignment for Decoupling Enrollment and Runtime Speaker Recognition Models
por: Gao, Chenyang, et al.
Publicado: (2024) -
Improving speaker verification robustness with synthetic emotional utterances
por: Koditala, Nikhil Kumar, et al.
Publicado: (2024)