Bangla-WhisperDiar: Fine-Tuning Whisper and PyAnnote for Bangla Long-Form Speech Recognition and Speaker Diarization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bhuiyan, Mohammed Aman, Adib, Md Sazzad Hossain, Bhuiyan, Samiul Basir, Chakraborty, Amit, Saswato, Aritra Islam, Dhrubo, Ahmed Faizul Haque, Khan, Mohammad Ashrafuzzaman
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913104145678336
author Bhuiyan, Mohammed Aman
Adib, Md Sazzad Hossain
Bhuiyan, Samiul Basir
Chakraborty, Amit
Saswato, Aritra Islam
Dhrubo, Ahmed Faizul Haque
Khan, Mohammad Ashrafuzzaman
author_facet Bhuiyan, Mohammed Aman
Adib, Md Sazzad Hossain
Bhuiyan, Samiul Basir
Chakraborty, Amit
Saswato, Aritra Islam
Dhrubo, Ahmed Faizul Haque
Khan, Mohammad Ashrafuzzaman
contents Automatic Speech Recognition (ASR) and speaker diarization in Bangla remain challenging due to long form recordings, diverse acoustic conditions, and significant speaker variability. This work addresses these two core tasks in Bangla spoken language understanding by developing robust systems for long form ASR and speaker diarization. For ASR (Problem 1), we fine tune the tugstugi bengaliai regional asr whisper medium model on a custom-curated dataset of approximately 15,000 chunked and aligned Bangla audio segments, employing full weight training with extensive data augmentation including noise injection, reverb simulation, echo, clipping distortion, and pitch/time perturbation. For speaker diarization (Problem 2), we fine-tune the pyannote/segmentation-3.0 model using PyTorch Lightning on the competition annotated diarization dataset, swapping the fine-tuned segmentation backbone into the pyannote/speaker-diarization-community-1 pipeline while retaining the pretrained speaker embedding and clustering components. Our ASR system achieves a Word Error Rate (WER) of 0.2441, while our diarization system achieves a Diarization Error Rate (DER) of 0.2392, both evaluated on the test set, demonstrating notable improvements over the respective pretrained baselines. We describe our complete pipeline, including data preprocessing, text normalization, audio augmentation, training strategies, inference optimization, and post-processing for both tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2605_08214
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Bangla-WhisperDiar: Fine-Tuning Whisper and PyAnnote for Bangla Long-Form Speech Recognition and Speaker Diarization
Bhuiyan, Mohammed Aman
Adib, Md Sazzad Hossain
Bhuiyan, Samiul Basir
Chakraborty, Amit
Saswato, Aritra Islam
Dhrubo, Ahmed Faizul Haque
Khan, Mohammad Ashrafuzzaman
Sound
Artificial Intelligence
Audio and Speech Processing
Automatic Speech Recognition (ASR) and speaker diarization in Bangla remain challenging due to long form recordings, diverse acoustic conditions, and significant speaker variability. This work addresses these two core tasks in Bangla spoken language understanding by developing robust systems for long form ASR and speaker diarization. For ASR (Problem 1), we fine tune the tugstugi bengaliai regional asr whisper medium model on a custom-curated dataset of approximately 15,000 chunked and aligned Bangla audio segments, employing full weight training with extensive data augmentation including noise injection, reverb simulation, echo, clipping distortion, and pitch/time perturbation. For speaker diarization (Problem 2), we fine-tune the pyannote/segmentation-3.0 model using PyTorch Lightning on the competition annotated diarization dataset, swapping the fine-tuned segmentation backbone into the pyannote/speaker-diarization-community-1 pipeline while retaining the pretrained speaker embedding and clustering components. Our ASR system achieves a Word Error Rate (WER) of 0.2441, while our diarization system achieves a Diarization Error Rate (DER) of 0.2392, both evaluated on the test set, demonstrating notable improvements over the respective pretrained baselines. We describe our complete pipeline, including data preprocessing, text normalization, audio augmentation, training strategies, inference optimization, and post-processing for both tasks.
title Bangla-WhisperDiar: Fine-Tuning Whisper and PyAnnote for Bangla Long-Form Speech Recognition and Speaker Diarization
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2605.08214