Pisets: A Robust Speech Recognition System for Lectures and Interviews
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911399409614848 |
|---|---|
| author | Bondarenko, Ivan Grebenkin, Daniil Sedukhin, Oleg Klementev, Mikhail Derunets, Roman Budneva, Lyudmila |
| author_facet | Bondarenko, Ivan Grebenkin, Daniil Sedukhin, Oleg Klementev, Mikhail Derunets, Roman Budneva, Lyudmila |
| contents | This work presents a speech-to-text system "Pisets" for scientists and journalists which is based on a three-component architecture aimed at improving speech recognition accuracy while minimizing errors and hallucinations associated with the Whisper model. The architecture comprises primary recognition using Wav2Vec2, false positive filtering via the Audio Spectrogram Transformer (AST), and final speech recognition through Whisper. The implementation of curriculum learning methods and the utilization of diverse Russian-language speech corpora significantly enhanced the system's effectiveness. Additionally, advanced uncertainty modeling techniques were introduced, contributing to further improvements in transcription quality. The proposed approaches ensure robust transcribing of long audio data across various acoustic conditions compared to WhisperX and the usual Whisper model. The source code of "Pisets" system is publicly available at GitHub: https://github.com/bond005/pisets. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_18415 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Pisets: A Robust Speech Recognition System for Lectures and Interviews Bondarenko, Ivan Grebenkin, Daniil Sedukhin, Oleg Klementev, Mikhail Derunets, Roman Budneva, Lyudmila Computation and Language Sound Audio and Speech Processing This work presents a speech-to-text system "Pisets" for scientists and journalists which is based on a three-component architecture aimed at improving speech recognition accuracy while minimizing errors and hallucinations associated with the Whisper model. The architecture comprises primary recognition using Wav2Vec2, false positive filtering via the Audio Spectrogram Transformer (AST), and final speech recognition through Whisper. The implementation of curriculum learning methods and the utilization of diverse Russian-language speech corpora significantly enhanced the system's effectiveness. Additionally, advanced uncertainty modeling techniques were introduced, contributing to further improvements in transcription quality. The proposed approaches ensure robust transcribing of long audio data across various acoustic conditions compared to WhisperX and the usual Whisper model. The source code of "Pisets" system is publicly available at GitHub: https://github.com/bond005/pisets. |
| title | Pisets: A Robust Speech Recognition System for Lectures and Interviews |
| topic | Computation and Language Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2601.18415 |