Vedavani: A Benchmark Corpus for ASR on Vedic Sanskrit Poetry

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kumar, Sujeet, Ray, Pretam, Beerukuri, Abhinay, Kamoji, Shrey, Jagadeeshan, Manoj Balaji, Goyal, Pawan
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908387792388096
author Kumar, Sujeet
Ray, Pretam
Beerukuri, Abhinay
Kamoji, Shrey
Jagadeeshan, Manoj Balaji
Goyal, Pawan
author_facet Kumar, Sujeet
Ray, Pretam
Beerukuri, Abhinay
Kamoji, Shrey
Jagadeeshan, Manoj Balaji
Goyal, Pawan
contents Sanskrit, an ancient language with a rich linguistic heritage, presents unique challenges for automatic speech recognition (ASR) due to its phonemic complexity and the phonetic transformations that occur at word junctures, similar to the connected speech found in natural conversations. Due to these complexities, there has been limited exploration of ASR in Sanskrit, particularly in the context of its poetic verses, which are characterized by intricate prosodic and rhythmic patterns. This gap in research raises the question: How can we develop an effective ASR system for Sanskrit, particularly one that captures the nuanced features of its poetic form? In this study, we introduce Vedavani, the first comprehensive ASR study focused on Sanskrit Vedic poetry. We present a 54-hour Sanskrit ASR dataset, consisting of 30,779 labelled audio samples from the Rig Veda and Atharva Veda. This dataset captures the precise prosodic and rhythmic features that define the language. We also benchmark the dataset on various state-of-the-art multilingual speech models.$^{1}$ Experimentation revealed that IndicWhisper performed the best among the SOTA models.
format Preprint
id arxiv_https___arxiv_org_abs_2506_00145
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Vedavani: A Benchmark Corpus for ASR on Vedic Sanskrit Poetry
Kumar, Sujeet
Ray, Pretam
Beerukuri, Abhinay
Kamoji, Shrey
Jagadeeshan, Manoj Balaji
Goyal, Pawan
Computation and Language
Sound
Audio and Speech Processing
Sanskrit, an ancient language with a rich linguistic heritage, presents unique challenges for automatic speech recognition (ASR) due to its phonemic complexity and the phonetic transformations that occur at word junctures, similar to the connected speech found in natural conversations. Due to these complexities, there has been limited exploration of ASR in Sanskrit, particularly in the context of its poetic verses, which are characterized by intricate prosodic and rhythmic patterns. This gap in research raises the question: How can we develop an effective ASR system for Sanskrit, particularly one that captures the nuanced features of its poetic form? In this study, we introduce Vedavani, the first comprehensive ASR study focused on Sanskrit Vedic poetry. We present a 54-hour Sanskrit ASR dataset, consisting of 30,779 labelled audio samples from the Rig Veda and Atharva Veda. This dataset captures the precise prosodic and rhythmic features that define the language. We also benchmark the dataset on various state-of-the-art multilingual speech models.$^{1}$ Experimentation revealed that IndicWhisper performed the best among the SOTA models.
title Vedavani: A Benchmark Corpus for ASR on Vedic Sanskrit Poetry
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2506.00145