Curriculum learning for self-supervised speaker verification
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Heo, Hee-Soo, Jung, Jee-weon, Kang, Jingu, Kwon, Youngki, Kim, You Jin, Lee, Bong-Jin, Chung, Joon Son |
|---|---|
| Format: | Preprint |
| Publié: |
2022
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Disentangled Representation Learning for Environment-agnostic Speaker Recognition
par: Nam, KiHyun, et autres
Publié: (2024)
par: Nam, KiHyun, et autres
Publié: (2024)
a-DCF: an architecture agnostic metric with application to spoofing-robust speaker verification
par: Shim, Hye-jin, et autres
Publié: (2024)
par: Shim, Hye-jin, et autres
Publié: (2024)
SEED: Speaker Embedding Enhancement Diffusion Model
par: Nam, KiHyun, et autres
Publié: (2025)
par: Nam, KiHyun, et autres
Publié: (2025)
Triage knowledge distillation for speaker verification
par: Kim, Ju-ho, et autres
Publié: (2026)
par: Kim, Ju-ho, et autres
Publié: (2026)
ESPnet-SPK: full pipeline speaker embedding toolkit with reproducible recipes, self-supervised front-ends, and off-the-shelf models
par: Jung, Jee-weon, et autres
Publié: (2024)
par: Jung, Jee-weon, et autres
Publié: (2024)
Optimization of DNN-based speaker verification model through efficient quantization technique
par: Hong, Yeona, et autres
Publié: (2024)
par: Hong, Yeona, et autres
Publié: (2024)
The VoxCeleb Speaker Recognition Challenge: A Retrospective
par: Huh, Jaesung, et autres
Publié: (2024)
par: Huh, Jaesung, et autres
Publié: (2024)
Text adaptation for speaker verification with speaker-text factorized embeddings
par: Yang, Yexin, et autres
Publié: (2025)
par: Yang, Yexin, et autres
Publié: (2025)
Improving Design of Input Condition Invariant Speech Enhancement
par: Zhang, Wangyou, et autres
Publié: (2024)
par: Zhang, Wangyou, et autres
Publié: (2024)
Token-based Attractors and Cross-attention in Spoof Diarization
par: Koo, Kyo-Won, et autres
Publié: (2025)
par: Koo, Kyo-Won, et autres
Publié: (2025)
Clustering-based hard negative sampling for supervised contrastive speaker verification
par: Masztalski, Piotr, et autres
Publié: (2025)
par: Masztalski, Piotr, et autres
Publié: (2025)
SCORE: Scaling audio generation using Standardized COmposite REwards
par: Jung, Jaemin, et autres
Publié: (2025)
par: Jung, Jaemin, et autres
Publié: (2025)
To what extent can ASV systems naturally defend against spoofing attacks?
par: Jung, Jee-weon, et autres
Publié: (2024)
par: Jung, Jee-weon, et autres
Publié: (2024)
Improving speaker verification robustness with synthetic emotional utterances
par: Koditala, Nikhil Kumar, et autres
Publié: (2024)
par: Koditala, Nikhil Kumar, et autres
Publié: (2024)
MF-PAM: Accurate Pitch Estimation through Periodicity Analysis and Multi-level Feature Fusion
par: Chung, Woo-Jin, et autres
Publié: (2023)
par: Chung, Woo-Jin, et autres
Publié: (2023)
Bridging the Gap between Audio and Text using Parallel-attention for User-defined Keyword Spotting
par: Kim, Youkyum, et autres
Publié: (2024)
par: Kim, Youkyum, et autres
Publié: (2024)
InfiniteAudio: Infinite-Length Audio Generation with Consistency
par: Jung, Chaeyoung, et autres
Publié: (2025)
par: Jung, Chaeyoung, et autres
Publié: (2025)
Tandem spoofing-robust automatic speaker verification based on time-domain embeddings
par: Weizman, Avishai, et autres
Publié: (2024)
par: Weizman, Avishai, et autres
Publié: (2024)
On the influence of language similarity in non-target speaker verification trials
par: Reuter, Paul M., et autres
Publié: (2025)
par: Reuter, Paul M., et autres
Publié: (2025)
Probing Cross-modal Information Hubs in Audio-Visual LLMs
par: Jung, Jihoo, et autres
Publié: (2026)
par: Jung, Jihoo, et autres
Publié: (2026)
FlowAVSE: Efficient Audio-Visual Speech Enhancement with Conditional Flow Matching
par: Jung, Chaeyoung, et autres
Publié: (2024)
par: Jung, Chaeyoung, et autres
Publié: (2024)
Gradient weighting for speaker verification in extremely low Signal-to-Noise Ratio
par: Ma, Yi, et autres
Publié: (2024)
par: Ma, Yi, et autres
Publié: (2024)
Improving fairness in speaker verification via Group-adapted Fusion Network
par: Shen, Hua, et autres
Publié: (2022)
par: Shen, Hua, et autres
Publié: (2022)
Beyond Performance Plateaus: A Comprehensive Study on Scalability in Speech Enhancement
par: Zhang, Wangyou, et autres
Publié: (2024)
par: Zhang, Wangyou, et autres
Publié: (2024)
A framework of text-dependent speaker verification for chinese numerical string corpus
par: Zheng, Litong, et autres
Publié: (2024)
par: Zheng, Litong, et autres
Publié: (2024)
LP-CFM: Perceptual Invariance-Aware Conditional Flow Matching for Speech Modeling
par: Kwak, Doyeop, et autres
Publié: (2025)
par: Kwak, Doyeop, et autres
Publié: (2025)
Improving curriculum learning for target speaker extraction with synthetic speakers
par: Liu, Yun, et autres
Publié: (2024)
par: Liu, Yun, et autres
Publié: (2024)
EDNet: A Versatile Speech Enhancement Framework with Gating Mamba Mechanism and Phase Shift-Invariant Training
par: Kwak, Doyeop, et autres
Publié: (2025)
par: Kwak, Doyeop, et autres
Publié: (2025)
MamTra: A Hybrid Mamba-Transformer Backbone for Speech Synthesis
par: Nguyen, Tan Dat, et autres
Publié: (2026)
par: Nguyen, Tan Dat, et autres
Publié: (2026)
Text-To-Speech Synthesis In The Wild
par: Jung, Jee-weon, et autres
Publié: (2024)
par: Jung, Jee-weon, et autres
Publié: (2024)
Beyond Silence: Bias Analysis through Loss and Asymmetric Approach in Audio Anti-Spoofing
par: Shim, Hye-jin, et autres
Publié: (2024)
par: Shim, Hye-jin, et autres
Publié: (2024)
Analyzing the relationships between pretraining language, phonetic, tonal, and speaker information in self-supervised speech models
par: Gubian, Michele, et autres
Publié: (2025)
par: Gubian, Michele, et autres
Publié: (2025)
UNMIXX: Untangling Highly Correlated Singing Voices Mixtures
par: Jung, Jihoo, et autres
Publié: (2026)
par: Jung, Jihoo, et autres
Publié: (2026)
Cinematic Audio Source Separation Using Visual Cues
par: Zhang, Kang, et autres
Publié: (2026)
par: Zhang, Kang, et autres
Publié: (2026)
Dub-S2ST: Textless Speech-to-Speech Translation for Seamless Dubbing
par: Choi, Jeongsoo, et autres
Publié: (2025)
par: Choi, Jeongsoo, et autres
Publié: (2025)
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis
par: Jung, Jaemin, et autres
Publié: (2024)
par: Jung, Jaemin, et autres
Publié: (2024)
Improving Audio Captioning Models with Fine-grained Audio Features, Text Embedding Supervision, and LLM Mix-up Augmentation
par: Wu, Shih-Lun, et autres
Publié: (2023)
par: Wu, Shih-Lun, et autres
Publié: (2023)
Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks
par: Maiti, Soumi, et autres
Publié: (2023)
par: Maiti, Soumi, et autres
Publié: (2023)
Quantifying the effect of speech pathology on automatic and human speaker verification
par: Halpern, Bence Mark, et autres
Publié: (2024)
par: Halpern, Bence Mark, et autres
Publié: (2024)
Team HYU ASML ROBOVOX SP Cup 2024 System Description
par: Choi, Jeong-Hwan, et autres
Publié: (2024)
par: Choi, Jeong-Hwan, et autres
Publié: (2024)
Documents similaires
-
Disentangled Representation Learning for Environment-agnostic Speaker Recognition
par: Nam, KiHyun, et autres
Publié: (2024) -
a-DCF: an architecture agnostic metric with application to spoofing-robust speaker verification
par: Shim, Hye-jin, et autres
Publié: (2024) -
SEED: Speaker Embedding Enhancement Diffusion Model
par: Nam, KiHyun, et autres
Publié: (2025) -
Triage knowledge distillation for speaker verification
par: Kim, Ju-ho, et autres
Publié: (2026) -
ESPnet-SPK: full pipeline speaker embedding toolkit with reproducible recipes, self-supervised front-ends, and off-the-shelf models
par: Jung, Jee-weon, et autres
Publié: (2024)