IITKGP-ABSP Submission to LRE22: Language Recognition in Low-Resource Settings

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dey, Spandan, Sahidullah, Md, Saha, Goutam
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916567123492864
author Dey, Spandan
Sahidullah, Md
Saha, Goutam
author_facet Dey, Spandan
Sahidullah, Md
Saha, Goutam
contents This is the detailed system description of the IITKGP-ABSP lab's submission to the NIST language recognition evaluation (LRE) 2022. The objective of this LRE (LRE22) is focused on recognizing 14 low-resourced African languages. Even though NIST has provided additional training and development data, we develop our systems with additional constraints of extreme low-resource. Our primary fixed-set submission ensures the usage of only the LRE 22 development data that contains the utterances of 14 target languages. We further restrict our system from using any pre-trained models for feature extraction or classifier fine-tuning. To address the issue of low-resource, our system relies on diverse audio augmentations followed by classifier fusions. Abiding by all the constraints, the proposed methods achieve an EER of 11.43% and cost metric of 0.41 in the LRE22 development set. For users with limited computational resources or limited storage/network capabilities, the proposed system will help achieve efficient LID performance.
format Preprint
id arxiv_https___arxiv_org_abs_2501_08616
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle IITKGP-ABSP Submission to LRE22: Language Recognition in Low-Resource Settings
Dey, Spandan
Sahidullah, Md
Saha, Goutam
Audio and Speech Processing
Sound
This is the detailed system description of the IITKGP-ABSP lab's submission to the NIST language recognition evaluation (LRE) 2022. The objective of this LRE (LRE22) is focused on recognizing 14 low-resourced African languages. Even though NIST has provided additional training and development data, we develop our systems with additional constraints of extreme low-resource. Our primary fixed-set submission ensures the usage of only the LRE 22 development data that contains the utterances of 14 target languages. We further restrict our system from using any pre-trained models for feature extraction or classifier fine-tuning. To address the issue of low-resource, our system relies on diverse audio augmentations followed by classifier fusions. Abiding by all the constraints, the proposed methods achieve an EER of 11.43% and cost metric of 0.41 in the LRE22 development set. For users with limited computational resources or limited storage/network capabilities, the proposed system will help achieve efficient LID performance.
title IITKGP-ABSP Submission to LRE22: Language Recognition in Low-Resource Settings
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2501.08616