What and When to Learn: CURriculum Ranking Loss for Large-Scale Speaker Verification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Baali, Massa, Bisht, Sarthak, Singh, Rita, Raj, Bhiksha
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908913375379456
author Baali, Massa
Bisht, Sarthak
Singh, Rita
Raj, Bhiksha
author_facet Baali, Massa
Bisht, Sarthak
Singh, Rita
Raj, Bhiksha
contents Speaker verification at large scale remains an open challenge as fixed-margin losses treat all samples equally regardless of quality. We hypothesize that mislabeled or degraded samples introduce noisy gradients that disrupt compact speaker manifolds. We propose Curry (CURriculum Ranking), an adaptive loss that estimates sample difficulty online via Sub-center ArcFace: confidence scores from dominant sub-center cosine similarity rank samples into easy, medium, and hard tiers using running batch statistics, without auxiliary annotations. Learnable weights guide the model from stable identity foundations through manifold refinement to boundary sharpening. To our knowledge, this is the largest-scale speaker verification system trained to date. Evaluated on VoxCeleb1-O, and SITW, Curry reduces EER by 86.8\% and 60.0\% over the Sub-center ArcFace baseline, establishing a new paradigm for robust speaker verification on imperfect large-scale data.
format Preprint
id arxiv_https___arxiv_org_abs_2603_24432
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle What and When to Learn: CURriculum Ranking Loss for Large-Scale Speaker Verification
Baali, Massa
Bisht, Sarthak
Singh, Rita
Raj, Bhiksha
Sound
Computation and Language
Speaker verification at large scale remains an open challenge as fixed-margin losses treat all samples equally regardless of quality. We hypothesize that mislabeled or degraded samples introduce noisy gradients that disrupt compact speaker manifolds. We propose Curry (CURriculum Ranking), an adaptive loss that estimates sample difficulty online via Sub-center ArcFace: confidence scores from dominant sub-center cosine similarity rank samples into easy, medium, and hard tiers using running batch statistics, without auxiliary annotations. Learnable weights guide the model from stable identity foundations through manifold refinement to boundary sharpening. To our knowledge, this is the largest-scale speaker verification system trained to date. Evaluated on VoxCeleb1-O, and SITW, Curry reduces EER by 86.8\% and 60.0\% over the Sub-center ArcFace baseline, establishing a new paradigm for robust speaker verification on imperfect large-scale data.
title What and When to Learn: CURriculum Ranking Loss for Large-Scale Speaker Verification
topic Sound
Computation and Language
url https://arxiv.org/abs/2603.24432