PERSONA: An Application for Emotion Recognition, Gender Recognition and Age Estimation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866907933524099072 |
|---|---|
| author | Koshal, Devyani Phukan, Orchid Chetia Jain, Sarthak Buduru, Arun Balaji Sharma, Rajesh |
| author_facet | Koshal, Devyani Phukan, Orchid Chetia Jain, Sarthak Buduru, Arun Balaji Sharma, Rajesh |
| contents | Emotion Recognition (ER), Gender Recognition (GR), and Age Estimation (AE) constitute paralinguistic tasks that rely not on the spoken content but primarily on speech characteristics such as pitch and tone. While previous research has made significant strides in developing models for each task individually, there has been comparatively less emphasis on concurrently learning these tasks, despite their inherent interconnectedness. As such in this demonstration, we present PERSONA, an application for predicting ER, GR, and AE with a single model in the backend. One notable point is we show that representations from speaker recognition pre-trained model (PTM) is better suited for such a multi-task learning format than the state-of-the-art (SOTA) self-supervised (SSL) PTM by carrying out a comparative study. Our methodology obviates the need for deploying separate models for each task and can potentially conserve resources and time during the training and deployment phases. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2406_06781 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | PERSONA: An Application for Emotion Recognition, Gender Recognition and Age Estimation Koshal, Devyani Phukan, Orchid Chetia Jain, Sarthak Buduru, Arun Balaji Sharma, Rajesh Audio and Speech Processing Sound Emotion Recognition (ER), Gender Recognition (GR), and Age Estimation (AE) constitute paralinguistic tasks that rely not on the spoken content but primarily on speech characteristics such as pitch and tone. While previous research has made significant strides in developing models for each task individually, there has been comparatively less emphasis on concurrently learning these tasks, despite their inherent interconnectedness. As such in this demonstration, we present PERSONA, an application for predicting ER, GR, and AE with a single model in the backend. One notable point is we show that representations from speaker recognition pre-trained model (PTM) is better suited for such a multi-task learning format than the state-of-the-art (SOTA) self-supervised (SSL) PTM by carrying out a comparative study. Our methodology obviates the need for deploying separate models for each task and can potentially conserve resources and time during the training and deployment phases. |
| title | PERSONA: An Application for Emotion Recognition, Gender Recognition and Age Estimation |
| topic | Audio and Speech Processing Sound |
| url | https://arxiv.org/abs/2406.06781 |