Exploring rhythm formant analysis for Indic language classification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gogoi, Parismita, Kalita, Sishir, Sarmah, Priyankoo, Prasanna, S. R Mahadeva
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917797384159232
author Gogoi, Parismita
Kalita, Sishir
Sarmah, Priyankoo
Prasanna, S. R Mahadeva
author_facet Gogoi, Parismita
Kalita, Sishir
Sarmah, Priyankoo
Prasanna, S. R Mahadeva
contents This paper reports a preliminary study on quantitative frequency domain rhythm cues for classifying five Indian languages: Bengali, Kannada, Malayalam, Marathi, and Tamil. We employ rhythm formant (R-formants) analysis, a technique introduced by Gibbon that utilizes low-frequency spectral analysis of amplitude modulation and frequency modulation envelopes to characterize speech rhythm. Various measures are computed from the LF spectrum, including R-formants, discrete cosine transform-based measures, and spectral measures. Results show that threshold-based and spectral features outperform directly computed R-formants. Temporal pattern of rhythm derived from LF spectrograms provides better language-discriminating cues. Combining all derived features we achieve an accuracy of 69.21% and a weighted F1 score of 69.18% in classifying the five languages. This study demonstrates the potential of RFA in characterizing speech rhythm for Indian language classification.
format Preprint
id arxiv_https___arxiv_org_abs_2410_05724
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Exploring rhythm formant analysis for Indic language classification
Gogoi, Parismita
Kalita, Sishir
Sarmah, Priyankoo
Prasanna, S. R Mahadeva
Audio and Speech Processing
Signal Processing
I.2.7
This paper reports a preliminary study on quantitative frequency domain rhythm cues for classifying five Indian languages: Bengali, Kannada, Malayalam, Marathi, and Tamil. We employ rhythm formant (R-formants) analysis, a technique introduced by Gibbon that utilizes low-frequency spectral analysis of amplitude modulation and frequency modulation envelopes to characterize speech rhythm. Various measures are computed from the LF spectrum, including R-formants, discrete cosine transform-based measures, and spectral measures. Results show that threshold-based and spectral features outperform directly computed R-formants. Temporal pattern of rhythm derived from LF spectrograms provides better language-discriminating cues. Combining all derived features we achieve an accuracy of 69.21% and a weighted F1 score of 69.18% in classifying the five languages. This study demonstrates the potential of RFA in characterizing speech rhythm for Indian language classification.
title Exploring rhythm formant analysis for Indic language classification
topic Audio and Speech Processing
Signal Processing
I.2.7
url https://arxiv.org/abs/2410.05724