Three-Class Emotion Classification for Audiovisual Scenes Based on Ensemble Learning Scheme

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xiong, Xiangrui, Zhou, Zhou, Nong, Guocai, Deng, Junlin, Wu, Ning
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915632792993792
author Xiong, Xiangrui
Zhou, Zhou
Nong, Guocai
Deng, Junlin
Wu, Ning
author_facet Xiong, Xiangrui
Zhou, Zhou
Nong, Guocai
Deng, Junlin
Wu, Ning
contents Emotion recognition plays a pivotal role in enhancing human-computer interaction, particularly in movie recommendation systems where understanding emotional content is essential. While multimodal approaches combining audio and video have demonstrated effectiveness, their reliance on high-performance graphical computing limits deployment on resource-constrained devices such as personal computers or home audiovisual systems. To address this limitation, this study proposes a novel audio-only ensemble learning framework capable of classifying movie scenes into three emotional categories: Good, Neutral, and Bad. The model integrates ten support vector machines and six neural networks within a stacking ensemble architecture to enhance classification performance. A tailored data preprocessing pipeline, including feature extraction, outlier handling, and feature engineering, is designed to optimize emotional information from audio inputs. Experiments on a simulated dataset achieve 67% accuracy, while a real-world dataset collected from 15 diverse films yields an impressive 86% accuracy. These results underscore the potential of audio-based, lightweight emotion recognition methods for broader consumer-level applications, offering both computational efficiency and robust classification capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2511_17926
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Three-Class Emotion Classification for Audiovisual Scenes Based on Ensemble Learning Scheme
Xiong, Xiangrui
Zhou, Zhou
Nong, Guocai
Deng, Junlin
Wu, Ning
Sound
Human-Computer Interaction
Emotion recognition plays a pivotal role in enhancing human-computer interaction, particularly in movie recommendation systems where understanding emotional content is essential. While multimodal approaches combining audio and video have demonstrated effectiveness, their reliance on high-performance graphical computing limits deployment on resource-constrained devices such as personal computers or home audiovisual systems. To address this limitation, this study proposes a novel audio-only ensemble learning framework capable of classifying movie scenes into three emotional categories: Good, Neutral, and Bad. The model integrates ten support vector machines and six neural networks within a stacking ensemble architecture to enhance classification performance. A tailored data preprocessing pipeline, including feature extraction, outlier handling, and feature engineering, is designed to optimize emotional information from audio inputs. Experiments on a simulated dataset achieve 67% accuracy, while a real-world dataset collected from 15 diverse films yields an impressive 86% accuracy. These results underscore the potential of audio-based, lightweight emotion recognition methods for broader consumer-level applications, offering both computational efficiency and robust classification capabilities.
title Three-Class Emotion Classification for Audiovisual Scenes Based on Ensemble Learning Scheme
topic Sound
Human-Computer Interaction
url https://arxiv.org/abs/2511.17926