Pisets: A Robust Speech Recognition System for Lectures and Interviews

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bondarenko, Ivan, Grebenkin, Daniil, Sedukhin, Oleg, Klementev, Mikhail, Derunets, Roman, Budneva, Lyudmila
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911399409614848
author Bondarenko, Ivan
Grebenkin, Daniil
Sedukhin, Oleg
Klementev, Mikhail
Derunets, Roman
Budneva, Lyudmila
author_facet Bondarenko, Ivan
Grebenkin, Daniil
Sedukhin, Oleg
Klementev, Mikhail
Derunets, Roman
Budneva, Lyudmila
contents This work presents a speech-to-text system "Pisets" for scientists and journalists which is based on a three-component architecture aimed at improving speech recognition accuracy while minimizing errors and hallucinations associated with the Whisper model. The architecture comprises primary recognition using Wav2Vec2, false positive filtering via the Audio Spectrogram Transformer (AST), and final speech recognition through Whisper. The implementation of curriculum learning methods and the utilization of diverse Russian-language speech corpora significantly enhanced the system's effectiveness. Additionally, advanced uncertainty modeling techniques were introduced, contributing to further improvements in transcription quality. The proposed approaches ensure robust transcribing of long audio data across various acoustic conditions compared to WhisperX and the usual Whisper model. The source code of "Pisets" system is publicly available at GitHub: https://github.com/bond005/pisets.
format Preprint
id arxiv_https___arxiv_org_abs_2601_18415
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Pisets: A Robust Speech Recognition System for Lectures and Interviews
Bondarenko, Ivan
Grebenkin, Daniil
Sedukhin, Oleg
Klementev, Mikhail
Derunets, Roman
Budneva, Lyudmila
Computation and Language
Sound
Audio and Speech Processing
This work presents a speech-to-text system "Pisets" for scientists and journalists which is based on a three-component architecture aimed at improving speech recognition accuracy while minimizing errors and hallucinations associated with the Whisper model. The architecture comprises primary recognition using Wav2Vec2, false positive filtering via the Audio Spectrogram Transformer (AST), and final speech recognition through Whisper. The implementation of curriculum learning methods and the utilization of diverse Russian-language speech corpora significantly enhanced the system's effectiveness. Additionally, advanced uncertainty modeling techniques were introduced, contributing to further improvements in transcription quality. The proposed approaches ensure robust transcribing of long audio data across various acoustic conditions compared to WhisperX and the usual Whisper model. The source code of "Pisets" system is publicly available at GitHub: https://github.com/bond005/pisets.
title Pisets: A Robust Speech Recognition System for Lectures and Interviews
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2601.18415