Agent-Based Modular Learning for Multimodal Emotion Recognition in Human-Agent Systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nepomnyaschiy, Matvey, Pereziabov, Oleg, Tliamov, Anvar, Mikhailov, Stanislav, Afanasyev, Ilya
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908707032399872
author Nepomnyaschiy, Matvey
Pereziabov, Oleg
Tliamov, Anvar
Mikhailov, Stanislav
Afanasyev, Ilya
author_facet Nepomnyaschiy, Matvey
Pereziabov, Oleg
Tliamov, Anvar
Mikhailov, Stanislav
Afanasyev, Ilya
contents Effective human-agent interaction (HAI) relies on accurate and adaptive perception of human emotional states. While multimodal deep learning models - leveraging facial expressions, speech, and textual cues - offer high accuracy in emotion recognition, their training and maintenance are often computationally intensive and inflexible to modality changes. In this work, we propose a novel multi-agent framework for training multimodal emotion recognition systems, where each modality encoder and the fusion classifier operate as autonomous agents coordinated by a central supervisor. This architecture enables modular integration of new modalities (e.g., audio features via emotion2vec), seamless replacement of outdated components, and reduced computational overhead during training. We demonstrate the feasibility of our approach through a proof-of-concept implementation supporting vision, audio, and text modalities, with the classifier serving as a shared decision-making agent. Our framework not only improves training efficiency but also contributes to the design of more flexible, scalable, and maintainable perception modules for embodied and virtual agents in HAI scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2512_10975
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Agent-Based Modular Learning for Multimodal Emotion Recognition in Human-Agent Systems
Nepomnyaschiy, Matvey
Pereziabov, Oleg
Tliamov, Anvar
Mikhailov, Stanislav
Afanasyev, Ilya
Machine Learning
Artificial Intelligence
Human-Computer Interaction
Multiagent Systems
Effective human-agent interaction (HAI) relies on accurate and adaptive perception of human emotional states. While multimodal deep learning models - leveraging facial expressions, speech, and textual cues - offer high accuracy in emotion recognition, their training and maintenance are often computationally intensive and inflexible to modality changes. In this work, we propose a novel multi-agent framework for training multimodal emotion recognition systems, where each modality encoder and the fusion classifier operate as autonomous agents coordinated by a central supervisor. This architecture enables modular integration of new modalities (e.g., audio features via emotion2vec), seamless replacement of outdated components, and reduced computational overhead during training. We demonstrate the feasibility of our approach through a proof-of-concept implementation supporting vision, audio, and text modalities, with the classifier serving as a shared decision-making agent. Our framework not only improves training efficiency but also contributes to the design of more flexible, scalable, and maintainable perception modules for embodied and virtual agents in HAI scenarios.
title Agent-Based Modular Learning for Multimodal Emotion Recognition in Human-Agent Systems
topic Machine Learning
Artificial Intelligence
Human-Computer Interaction
Multiagent Systems
url https://arxiv.org/abs/2512.10975