MERaLiON-SER: Robust Speech Emotion Recognition Model for English and SEA Languages

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Sailor, Hardik B., Ti, Aw Ai, Nancy, Chen Fang Yih, Lay, Chiu Ying, Yang, Ding, Yingxu, He, Ridong, Jiang, Jingtao, Li, Jingyi, Liao, Zhuohan, Liu, Yanfeng, Lu, Yi, Ma, Gupta, Manas, Shahrin, Muhammad Huzaifah Bin Md, Johan, Nabilah Binte Md, Lertcheva, Nattadaporn, Chunlei, Pan, Duc, Pham Minh, Subaidi, Siti Maryam Binte Ahmad, Salleh, Siti Umairah Binte Mohammad, Shuo, Sun, Vangani, Tarun Kumar, Qiongqiong, Wang, Lewis, Won Cheng Yi, Jeremy, Wong Heng Meng, Jinyang, Wu, Huayun, Zhang, Longyin, Zhang, Xunlong, Zou
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917075244548096
author Sailor, Hardik B.
Ti, Aw Ai
Nancy, Chen Fang Yih
Lay, Chiu Ying
Yang, Ding
Yingxu, He
Ridong, Jiang
Jingtao, Li
Jingyi, Liao
Zhuohan, Liu
Yanfeng, Lu
Yi, Ma
Gupta, Manas
Shahrin, Muhammad Huzaifah Bin Md
Johan, Nabilah Binte Md
Lertcheva, Nattadaporn
Chunlei, Pan
Duc, Pham Minh
Subaidi, Siti Maryam Binte Ahmad
Salleh, Siti Umairah Binte Mohammad
Shuo, Sun
Vangani, Tarun Kumar
Qiongqiong, Wang
Lewis, Won Cheng Yi
Jeremy, Wong Heng Meng
Jinyang, Wu
Huayun, Zhang
Longyin, Zhang
Xunlong, Zou
author_facet Sailor, Hardik B.
Ti, Aw Ai
Nancy, Chen Fang Yih
Lay, Chiu Ying
Yang, Ding
Yingxu, He
Ridong, Jiang
Jingtao, Li
Jingyi, Liao
Zhuohan, Liu
Yanfeng, Lu
Yi, Ma
Gupta, Manas
Shahrin, Muhammad Huzaifah Bin Md
Johan, Nabilah Binte Md
Lertcheva, Nattadaporn
Chunlei, Pan
Duc, Pham Minh
Subaidi, Siti Maryam Binte Ahmad
Salleh, Siti Umairah Binte Mohammad
Shuo, Sun
Vangani, Tarun Kumar
Qiongqiong, Wang
Lewis, Won Cheng Yi
Jeremy, Wong Heng Meng
Jinyang, Wu
Huayun, Zhang
Longyin, Zhang
Xunlong, Zou
contents We present MERaLiON-SER, a robust speech emotion recognition model designed for English and Southeast Asian languages. The model is trained using a hybrid objective combining weighted categorical cross-entropy and Concordance Correlation Coefficient (CCC) losses for joint discrete and dimensional emotion modelling. This dual approach enables the model to capture both the distinct categories of emotion (like happy or angry) and the fine-grained, such as arousal (intensity), valence (positivity/negativity), and dominance (sense of control), leading to a more comprehensive and robust representation of human affect. Extensive evaluations across multilingual Singaporean languages (English, Chinese, Malay, and Tamil ) and other public benchmarks show that MERaLiON-SER consistently surpasses both open-source speech encoders and large Audio-LLMs. These results underscore the importance of specialised speech-only models for accurate paralinguistic understanding and cross-lingual generalisation. Furthermore, the proposed framework provides a foundation for integrating emotion-aware perception into future agentic audio systems, enabling more empathetic and contextually adaptive multimodal reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2511_04914
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MERaLiON-SER: Robust Speech Emotion Recognition Model for English and SEA Languages
Sailor, Hardik B.
Ti, Aw Ai
Nancy, Chen Fang Yih
Lay, Chiu Ying
Yang, Ding
Yingxu, He
Ridong, Jiang
Jingtao, Li
Jingyi, Liao
Zhuohan, Liu
Yanfeng, Lu
Yi, Ma
Gupta, Manas
Shahrin, Muhammad Huzaifah Bin Md
Johan, Nabilah Binte Md
Lertcheva, Nattadaporn
Chunlei, Pan
Duc, Pham Minh
Subaidi, Siti Maryam Binte Ahmad
Salleh, Siti Umairah Binte Mohammad
Shuo, Sun
Vangani, Tarun Kumar
Qiongqiong, Wang
Lewis, Won Cheng Yi
Jeremy, Wong Heng Meng
Jinyang, Wu
Huayun, Zhang
Longyin, Zhang
Xunlong, Zou
Sound
Artificial Intelligence
We present MERaLiON-SER, a robust speech emotion recognition model designed for English and Southeast Asian languages. The model is trained using a hybrid objective combining weighted categorical cross-entropy and Concordance Correlation Coefficient (CCC) losses for joint discrete and dimensional emotion modelling. This dual approach enables the model to capture both the distinct categories of emotion (like happy or angry) and the fine-grained, such as arousal (intensity), valence (positivity/negativity), and dominance (sense of control), leading to a more comprehensive and robust representation of human affect. Extensive evaluations across multilingual Singaporean languages (English, Chinese, Malay, and Tamil ) and other public benchmarks show that MERaLiON-SER consistently surpasses both open-source speech encoders and large Audio-LLMs. These results underscore the importance of specialised speech-only models for accurate paralinguistic understanding and cross-lingual generalisation. Furthermore, the proposed framework provides a foundation for integrating emotion-aware perception into future agentic audio systems, enabling more empathetic and contextually adaptive multimodal reasoning.
title MERaLiON-SER: Robust Speech Emotion Recognition Model for English and SEA Languages
topic Sound
Artificial Intelligence
url https://arxiv.org/abs/2511.04914