VoxBlink2: A 100K+ Speaker Recognition Corpus and the Open-Set Speaker-Identification Benchmark
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lin, Yuke, Cheng, Ming, Zhang, Fulin, Gao, Yingying, Zhang, Shilei, Li, Ming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CEC: A Noisy Label Detection Method for Speaker Recognition
von: Shen, Yao, et al.
Veröffentlicht: (2024)
von: Shen, Yao, et al.
Veröffentlicht: (2024)
Sequence-to-Sequence Neural Diarization with Automatic Speaker Detection and Representation
von: Cheng, Ming, et al.
Veröffentlicht: (2024)
von: Cheng, Ming, et al.
Veröffentlicht: (2024)
The DKU System for Multi-Speaker Automatic Speech Recognition in MLC-SLM Challenge
von: Lin, Yuke, et al.
Veröffentlicht: (2025)
von: Lin, Yuke, et al.
Veröffentlicht: (2025)
Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models
von: Lin, Yuke, et al.
Veröffentlicht: (2025)
von: Lin, Yuke, et al.
Veröffentlicht: (2025)
The Database and Benchmark for the Source Speaker Tracing Challenge 2024
von: Li, Ze, et al.
Veröffentlicht: (2024)
von: Li, Ze, et al.
Veröffentlicht: (2024)
Multi-Input Multi-Output Target-Speaker Voice Activity Detection For Unified, Flexible, and Robust Audio-Visual Speaker Diarization
von: Cheng, Ming, et al.
Veröffentlicht: (2024)
von: Cheng, Ming, et al.
Veröffentlicht: (2024)
Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data
von: Liu, Yun, et al.
Veröffentlicht: (2024)
von: Liu, Yun, et al.
Veröffentlicht: (2024)
The VoxCeleb Speaker Recognition Challenge: A Retrospective
von: Huh, Jaesung, et al.
Veröffentlicht: (2024)
von: Huh, Jaesung, et al.
Veröffentlicht: (2024)
Enhancing Open-Set Speaker Identification through Rapid Tuning with Speaker Reciprocal Points and Negative Sample
von: Chen, Zhiyong, et al.
Veröffentlicht: (2024)
von: Chen, Zhiyong, et al.
Veröffentlicht: (2024)
KunquDB: An Attempt for Speaker Verification in the Chinese Opera Scenario
von: Zhou, Huali, et al.
Veröffentlicht: (2024)
von: Zhou, Huali, et al.
Veröffentlicht: (2024)
USEF-TSE: Universal Speaker Embedding Free Target Speaker Extraction
von: Zeng, Bang, et al.
Veröffentlicht: (2024)
von: Zeng, Bang, et al.
Veröffentlicht: (2024)
Study on Inter and Intra Speaker Variability in Speaker Recognition
von: Okhotnikov, Anton, et al.
Veröffentlicht: (2024)
von: Okhotnikov, Anton, et al.
Veröffentlicht: (2024)
Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits
von: Feng, Tiantian, et al.
Veröffentlicht: (2025)
von: Feng, Tiantian, et al.
Veröffentlicht: (2025)
Universal Speaker Embedding Free Target Speaker Extraction and Personal Voice Activity Detection
von: Zeng, Bang, et al.
Veröffentlicht: (2025)
von: Zeng, Bang, et al.
Veröffentlicht: (2025)
oboVox Far Field Speaker Recognition: A Novel Data Augmentation Approach with Pretrained Models
von: Dip, Muhammad Sudipto Siam, et al.
Veröffentlicht: (2024)
von: Dip, Muhammad Sudipto Siam, et al.
Veröffentlicht: (2024)
Emotion Recognition in Multi-Speaker Conversations through Speaker Identification, Knowledge Distillation, and Hierarchical Fusion
von: Li, Xiao, et al.
Veröffentlicht: (2025)
von: Li, Xiao, et al.
Veröffentlicht: (2025)
MFSN: Multi-perspective Fusion Search Network For Pre-training Knowledge in Speech Emotion Recognition
von: Sun, Haiyang, et al.
Veröffentlicht: (2023)
von: Sun, Haiyang, et al.
Veröffentlicht: (2023)
Can We Really Repurpose Multi-Speaker ASR Corpus for Speaker Diarization?
von: Horiguchi, Shota, et al.
Veröffentlicht: (2025)
von: Horiguchi, Shota, et al.
Veröffentlicht: (2025)
Enhancing Speaker Verification with w2v-BERT 2.0 and Knowledge Distillation guided Structured Pruning
von: Li, Ze, et al.
Veröffentlicht: (2025)
von: Li, Ze, et al.
Veröffentlicht: (2025)
Rhythm Features for Speaker Identification
von: Mehlman, Nick, et al.
Veröffentlicht: (2025)
von: Mehlman, Nick, et al.
Veröffentlicht: (2025)
SpeakerRPL v2: Robust Open-set Speaker Identification through Enhanced Few-shot Foundation Tuning and Model Fusion
von: Chen, Zhiyong, et al.
Veröffentlicht: (2026)
von: Chen, Zhiyong, et al.
Veröffentlicht: (2026)
3D-Speaker-Toolkit: An Open-Source Toolkit for Multimodal Speaker Verification and Diarization
von: Chen, Yafeng, et al.
Veröffentlicht: (2024)
von: Chen, Yafeng, et al.
Veröffentlicht: (2024)
Pretraining Multi-Speaker Identification for Neural Speaker Diarization
von: Horiguchi, Shota, et al.
Veröffentlicht: (2025)
von: Horiguchi, Shota, et al.
Veröffentlicht: (2025)
A Comprehensive Investigation on Speaker Augmentation for Speaker Recognition
von: Zhou, Zhenyu, et al.
Veröffentlicht: (2024)
von: Zhou, Zhenyu, et al.
Veröffentlicht: (2024)
VoxVietnam: a Large-Scale Multi-Genre Dataset for Vietnamese Speaker Recognition
von: Vu, Hoang Long, et al.
Veröffentlicht: (2024)
von: Vu, Hoang Long, et al.
Veröffentlicht: (2024)
Reshape Dimensions Network for Speaker Recognition
von: Yakovlev, Ivan, et al.
Veröffentlicht: (2024)
von: Yakovlev, Ivan, et al.
Veröffentlicht: (2024)
VoxGenesis: Unsupervised Discovery of Latent Speaker Manifold for Speech Synthesis
von: Lin, Weiwei, et al.
Veröffentlicht: (2024)
von: Lin, Weiwei, et al.
Veröffentlicht: (2024)
B-GRPO: Unsupervised Speech Emotion Recognition based on Batched-Group Relative Policy Optimization
von: Gao, Yingying, et al.
Veröffentlicht: (2026)
von: Gao, Yingying, et al.
Veröffentlicht: (2026)
Target Speaker Lipreading by Audio-Visual Self-Distillation Pretraining and Speaker Adaptation
von: Zhang, Jing-Xuan, et al.
Veröffentlicht: (2025)
von: Zhang, Jing-Xuan, et al.
Veröffentlicht: (2025)
Leveraging ASR Pretrained Conformers for Speaker Verification through Transfer Learning and Knowledge Distillation
von: Cai, Danwei, et al.
Veröffentlicht: (2023)
von: Cai, Danwei, et al.
Veröffentlicht: (2023)
Multi-View Based Audio Visual Target Speaker Extraction
von: Yang, Peijun, et al.
Veröffentlicht: (2026)
von: Yang, Peijun, et al.
Veröffentlicht: (2026)
A Toolkit for Joint Speaker Diarization and Identification with Application to Speaker-Attributed ASR
von: Morrone, Giovanni, et al.
Veröffentlicht: (2024)
von: Morrone, Giovanni, et al.
Veröffentlicht: (2024)
Quantifying and Reducing Speaker Heterogeneity within the Common Voice Corpus for Phonetic Analysis
von: Zhang, Miao, et al.
Veröffentlicht: (2025)
von: Zhang, Miao, et al.
Veröffentlicht: (2025)
Speaker Contrastive Learning for Source Speaker Tracing
von: Wang, Qing, et al.
Veröffentlicht: (2024)
von: Wang, Qing, et al.
Veröffentlicht: (2024)
Language-Invariant Multilingual Speaker Verification for the TidyVoice 2026 Challenge
von: Li, Ze, et al.
Veröffentlicht: (2026)
von: Li, Ze, et al.
Veröffentlicht: (2026)
openFEAT: Improving Speaker Identification by Open-set Few-shot Embedding Adaptation with Transformer
von: C, Kishan K, et al.
Veröffentlicht: (2022)
von: C, Kishan K, et al.
Veröffentlicht: (2022)
Multi-Channel Multi-Speaker ASR Using Target Speaker's Solo Segment
von: Shao, Yiwen, et al.
Veröffentlicht: (2024)
von: Shao, Yiwen, et al.
Veröffentlicht: (2024)
Emotional Styles Hide in Deep Speaker Embeddings: Disentangle Deep Speaker Embeddings for Speaker Clustering
von: Lin, Chaohao, et al.
Veröffentlicht: (2025)
von: Lin, Chaohao, et al.
Veröffentlicht: (2025)
Voice Conversion Augmentation for Speaker Recognition on Defective Datasets
von: Tao, Ruijie, et al.
Veröffentlicht: (2024)
von: Tao, Ruijie, et al.
Veröffentlicht: (2024)
Multi-Level Speaker Representation for Target Speaker Extraction
von: Zhang, Ke, et al.
Veröffentlicht: (2024)
von: Zhang, Ke, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
CEC: A Noisy Label Detection Method for Speaker Recognition
von: Shen, Yao, et al.
Veröffentlicht: (2024) -
Sequence-to-Sequence Neural Diarization with Automatic Speaker Detection and Representation
von: Cheng, Ming, et al.
Veröffentlicht: (2024) -
The DKU System for Multi-Speaker Automatic Speech Recognition in MLC-SLM Challenge
von: Lin, Yuke, et al.
Veröffentlicht: (2025) -
Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models
von: Lin, Yuke, et al.
Veröffentlicht: (2025) -
The Database and Benchmark for the Source Speaker Tracing Challenge 2024
von: Li, Ze, et al.
Veröffentlicht: (2024)