Investigating Polyglot Speech Foundation Models for Learning Collective Emotion from Crowds

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Phukan, Orchid Chetia, Girish, Akhtar, Mohd Mujtaba, Nayak, Panchal, Mallick, Priyabrata, Behera, Swarup Ranjan, Bhagath, Parabattina, Reddy, Pailla Balakrishna, Buduru, Arun Balaji
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912595227705344
author Phukan, Orchid Chetia
Girish
Akhtar, Mohd Mujtaba
Nayak, Panchal
Mallick, Priyabrata
Behera, Swarup Ranjan
Bhagath, Parabattina
Reddy, Pailla Balakrishna
Buduru, Arun Balaji
author_facet Phukan, Orchid Chetia
Girish
Akhtar, Mohd Mujtaba
Nayak, Panchal
Mallick, Priyabrata
Behera, Swarup Ranjan
Bhagath, Parabattina
Reddy, Pailla Balakrishna
Buduru, Arun Balaji
contents This paper investigates the polyglot (multilingual) speech foundation models (SFMs) for Crowd Emotion Recognition (CER). We hypothesize that polyglot SFMs, pre-trained on diverse languages, accents, and speech patterns, are particularly adept at navigating the noisy and complex acoustic environments characteristic of crowd settings, thereby offering a significant advantage for CER. To substantiate this, we perform a comprehensive analysis, comparing polyglot, monolingual, and speaker recognition SFMs through extensive experiments on a benchmark CER dataset across varying audio durations (1 sec, 500 ms, and 250 ms). The results consistently demonstrate the superiority of polyglot SFMs, outperforming their counterparts across all audio lengths and excelling even with extremely short-duration inputs. These findings pave the way for adaptation of SFMs in setting up new benchmarks for CER.
format Preprint
id arxiv_https___arxiv_org_abs_2509_16329
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Investigating Polyglot Speech Foundation Models for Learning Collective Emotion from Crowds
Phukan, Orchid Chetia
Girish
Akhtar, Mohd Mujtaba
Nayak, Panchal
Mallick, Priyabrata
Behera, Swarup Ranjan
Bhagath, Parabattina
Reddy, Pailla Balakrishna
Buduru, Arun Balaji
Audio and Speech Processing
Sound
This paper investigates the polyglot (multilingual) speech foundation models (SFMs) for Crowd Emotion Recognition (CER). We hypothesize that polyglot SFMs, pre-trained on diverse languages, accents, and speech patterns, are particularly adept at navigating the noisy and complex acoustic environments characteristic of crowd settings, thereby offering a significant advantage for CER. To substantiate this, we perform a comprehensive analysis, comparing polyglot, monolingual, and speaker recognition SFMs through extensive experiments on a benchmark CER dataset across varying audio durations (1 sec, 500 ms, and 250 ms). The results consistently demonstrate the superiority of polyglot SFMs, outperforming their counterparts across all audio lengths and excelling even with extremely short-duration inputs. These findings pave the way for adaptation of SFMs in setting up new benchmarks for CER.
title Investigating Polyglot Speech Foundation Models for Learning Collective Emotion from Crowds
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2509.16329