BiRQ: Bi-Level Self-Labeling Random Quantization for Self-Supervised Speech Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Liuyuan, Cui, Xiaodong, Kingsbury, Brian, Chen, Tianyi, Chen, Lisha
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908546458714112
author Jiang, Liuyuan
Cui, Xiaodong
Kingsbury, Brian
Chen, Tianyi
Chen, Lisha
author_facet Jiang, Liuyuan
Cui, Xiaodong
Kingsbury, Brian
Chen, Tianyi
Chen, Lisha
contents Speech is a rich signal, and labeled audio-text pairs are costly, making self-supervised learning essential for scalable representation learning. A core challenge in speech SSL is generating pseudo-labels that are both informative and efficient: strong labels, such as those used in HuBERT, improve downstream performance but rely on external encoders and multi-stage pipelines, while efficient methods like BEST-RQ achieve simplicity at the cost of weaker labels. We propose BiRQ, a bilevel SSL framework that combines the efficiency of BEST-RQ with the refinement benefits of HuBERT-style label enhancement. The key idea is to reuse part of the model itself as a pseudo-label generator: intermediate representations are discretized by a random-projection quantizer to produce enhanced labels, while anchoring labels derived directly from the raw input stabilize training and prevent collapse. Training is formulated as an efficient first-order bilevel optimization problem, solved end-to-end with differentiable Gumbel-softmax selection. This design eliminates the need for external label encoders, reduces memory cost, and enables iterative label refinement in an end-to-end fashion. BiRQ consistently improves over BEST-RQ while maintaining low complexity and computational efficiency. We validate our method on various datasets, including 960-hour LibriSpeech, 150-hour AMI meetings and 5,000-hour YODAS, demonstrating consistent gains over BEST-RQ.
format Preprint
id arxiv_https___arxiv_org_abs_2509_15430
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle BiRQ: Bi-Level Self-Labeling Random Quantization for Self-Supervised Speech Recognition
Jiang, Liuyuan
Cui, Xiaodong
Kingsbury, Brian
Chen, Tianyi
Chen, Lisha
Computation and Language
Sound
Audio and Speech Processing
Speech is a rich signal, and labeled audio-text pairs are costly, making self-supervised learning essential for scalable representation learning. A core challenge in speech SSL is generating pseudo-labels that are both informative and efficient: strong labels, such as those used in HuBERT, improve downstream performance but rely on external encoders and multi-stage pipelines, while efficient methods like BEST-RQ achieve simplicity at the cost of weaker labels. We propose BiRQ, a bilevel SSL framework that combines the efficiency of BEST-RQ with the refinement benefits of HuBERT-style label enhancement. The key idea is to reuse part of the model itself as a pseudo-label generator: intermediate representations are discretized by a random-projection quantizer to produce enhanced labels, while anchoring labels derived directly from the raw input stabilize training and prevent collapse. Training is formulated as an efficient first-order bilevel optimization problem, solved end-to-end with differentiable Gumbel-softmax selection. This design eliminates the need for external label encoders, reduces memory cost, and enables iterative label refinement in an end-to-end fashion. BiRQ consistently improves over BEST-RQ while maintaining low complexity and computational efficiency. We validate our method on various datasets, including 960-hour LibriSpeech, 150-hour AMI meetings and 5,000-hour YODAS, demonstrating consistent gains over BEST-RQ.
title BiRQ: Bi-Level Self-Labeling Random Quantization for Self-Supervised Speech Recognition
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2509.15430