PERSONA: An Application for Emotion Recognition, Gender Recognition and Age Estimation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Koshal, Devyani, Phukan, Orchid Chetia, Jain, Sarthak, Buduru, Arun Balaji, Sharma, Rajesh
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866907933524099072
author Koshal, Devyani
Phukan, Orchid Chetia
Jain, Sarthak
Buduru, Arun Balaji
Sharma, Rajesh
author_facet Koshal, Devyani
Phukan, Orchid Chetia
Jain, Sarthak
Buduru, Arun Balaji
Sharma, Rajesh
contents Emotion Recognition (ER), Gender Recognition (GR), and Age Estimation (AE) constitute paralinguistic tasks that rely not on the spoken content but primarily on speech characteristics such as pitch and tone. While previous research has made significant strides in developing models for each task individually, there has been comparatively less emphasis on concurrently learning these tasks, despite their inherent interconnectedness. As such in this demonstration, we present PERSONA, an application for predicting ER, GR, and AE with a single model in the backend. One notable point is we show that representations from speaker recognition pre-trained model (PTM) is better suited for such a multi-task learning format than the state-of-the-art (SOTA) self-supervised (SSL) PTM by carrying out a comparative study. Our methodology obviates the need for deploying separate models for each task and can potentially conserve resources and time during the training and deployment phases.
format Preprint
id arxiv_https___arxiv_org_abs_2406_06781
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle PERSONA: An Application for Emotion Recognition, Gender Recognition and Age Estimation
Koshal, Devyani
Phukan, Orchid Chetia
Jain, Sarthak
Buduru, Arun Balaji
Sharma, Rajesh
Audio and Speech Processing
Sound
Emotion Recognition (ER), Gender Recognition (GR), and Age Estimation (AE) constitute paralinguistic tasks that rely not on the spoken content but primarily on speech characteristics such as pitch and tone. While previous research has made significant strides in developing models for each task individually, there has been comparatively less emphasis on concurrently learning these tasks, despite their inherent interconnectedness. As such in this demonstration, we present PERSONA, an application for predicting ER, GR, and AE with a single model in the backend. One notable point is we show that representations from speaker recognition pre-trained model (PTM) is better suited for such a multi-task learning format than the state-of-the-art (SOTA) self-supervised (SSL) PTM by carrying out a comparative study. Our methodology obviates the need for deploying separate models for each task and can potentially conserve resources and time during the training and deployment phases.
title PERSONA: An Application for Emotion Recognition, Gender Recognition and Age Estimation
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2406.06781