Saved in:
Bibliographic Details
Main Authors: Fang, Yiyang, Huang, Wenke, Fu, Pei, Yang, Yihao, Su, Kehua, Luo, Zhenbo, Luan, Jian, Ye, Mang
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2602.23802
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917299289587712
author Fang, Yiyang
Huang, Wenke
Fu, Pei
Yang, Yihao
Su, Kehua
Luo, Zhenbo
Luan, Jian
Ye, Mang
author_facet Fang, Yiyang
Huang, Wenke
Fu, Pei
Yang, Yihao
Su, Kehua
Luo, Zhenbo
Luan, Jian
Ye, Mang
contents Multimodal Large Language Models (MLLMs) have shown remarkable progress in visual reasoning and understanding tasks but still struggle to capture the complexity and subjectivity of human emotions. Existing approaches based on supervised fine-tuning often suffer from limited generalization and poor interpretability, while reinforcement learning methods such as Group Relative Policy Optimization fail to align with the intrinsic characteristics of emotional cognition. To address these challenges, we propose Reflective Reinforcement Learning for Emotional Reasoning (EMO-R3), a framework designed to enhance the emotional reasoning ability of MLLMs. Specifically, we introduce Structured Emotional Thinking to guide the model to perform step-by-step emotional reasoning in a structured and interpretable manner, and design a Reflective Emotional Reward that enables the model to re-evaluate its reasoning based on visual-text consistency and emotional coherence. Extensive experiments demonstrate that EMO-R3 significantly improves both the interpretability and emotional intelligence of MLLMs, achieving superior performance across multiple visual emotional understanding benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2602_23802
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models
Fang, Yiyang
Huang, Wenke
Fu, Pei
Yang, Yihao
Su, Kehua
Luo, Zhenbo
Luan, Jian
Ye, Mang
Artificial Intelligence
Computer Vision and Pattern Recognition
Multimodal Large Language Models (MLLMs) have shown remarkable progress in visual reasoning and understanding tasks but still struggle to capture the complexity and subjectivity of human emotions. Existing approaches based on supervised fine-tuning often suffer from limited generalization and poor interpretability, while reinforcement learning methods such as Group Relative Policy Optimization fail to align with the intrinsic characteristics of emotional cognition. To address these challenges, we propose Reflective Reinforcement Learning for Emotional Reasoning (EMO-R3), a framework designed to enhance the emotional reasoning ability of MLLMs. Specifically, we introduce Structured Emotional Thinking to guide the model to perform step-by-step emotional reasoning in a structured and interpretable manner, and design a Reflective Emotional Reward that enables the model to re-evaluate its reasoning based on visual-text consistency and emotional coherence. Extensive experiments demonstrate that EMO-R3 significantly improves both the interpretability and emotional intelligence of MLLMs, achieving superior performance across multiple visual emotional understanding benchmarks.
title EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models
topic Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.23802