The TCG CREST -- RKMVERI Submission for the NCIIPC Startup India AI Grand Challenge

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Raghav, Nikhil, Banerjee, Arnab, Chakraborty, Janojit, Gupta, Avisek, Punyeshwarananda, Swami, Sahidullah, Md
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909957505417216
author Raghav, Nikhil
Banerjee, Arnab
Chakraborty, Janojit
Gupta, Avisek
Punyeshwarananda, Swami
Sahidullah, Md
author_facet Raghav, Nikhil
Banerjee, Arnab
Chakraborty, Janojit
Gupta, Avisek
Punyeshwarananda, Swami
Sahidullah, Md
contents In this report, we summarize the integrated multilingual audio processing pipeline developed by our team for the inaugural NCIIPC Startup India AI GRAND CHALLENGE, addressing Problem Statement 06: Language-Agnostic Speaker Identification and Diarisation, and subsequent Transcription and Translation System. Our primary focus was on advancing speaker diarization, a critical component for multilingual and code-mixed scenarios. The main intent of this work was to study the real-world applicability of our in-house speaker diarization (SD) systems. To this end, we investigated a robust voice activity detection (VAD) technique and fine-tuned speaker embedding models for improved speaker identification in low-resource settings. We leveraged our own recently proposed multi-kernel consensus spectral clustering framework, which substantially improved the diarization performance across all recordings in the training corpus provided by the organizers. Complementary modules for speaker and language identification, automatic speech recognition (ASR), and neural machine translation were integrated in the pipeline. Post-processing refinements further improved system robustness.
format Preprint
id arxiv_https___arxiv_org_abs_2512_11009
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The TCG CREST -- RKMVERI Submission for the NCIIPC Startup India AI Grand Challenge
Raghav, Nikhil
Banerjee, Arnab
Chakraborty, Janojit
Gupta, Avisek
Punyeshwarananda, Swami
Sahidullah, Md
Sound
In this report, we summarize the integrated multilingual audio processing pipeline developed by our team for the inaugural NCIIPC Startup India AI GRAND CHALLENGE, addressing Problem Statement 06: Language-Agnostic Speaker Identification and Diarisation, and subsequent Transcription and Translation System. Our primary focus was on advancing speaker diarization, a critical component for multilingual and code-mixed scenarios. The main intent of this work was to study the real-world applicability of our in-house speaker diarization (SD) systems. To this end, we investigated a robust voice activity detection (VAD) technique and fine-tuned speaker embedding models for improved speaker identification in low-resource settings. We leveraged our own recently proposed multi-kernel consensus spectral clustering framework, which substantially improved the diarization performance across all recordings in the training corpus provided by the organizers. Complementary modules for speaker and language identification, automatic speech recognition (ASR), and neural machine translation were integrated in the pipeline. Post-processing refinements further improved system robustness.
title The TCG CREST -- RKMVERI Submission for the NCIIPC Startup India AI Grand Challenge
topic Sound
url https://arxiv.org/abs/2512.11009