A cross-modal network for facial expression recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tian, Chunwei, Xie, Jingyuan, Zhang, Qi, Li, Chao, Zuo, Wangmeng, Zhang, Shichao
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918485880209408
author Tian, Chunwei
Xie, Jingyuan
Zhang, Qi
Li, Chao
Zuo, Wangmeng
Zhang, Shichao
author_facet Tian, Chunwei
Xie, Jingyuan
Zhang, Qi
Li, Chao
Zuo, Wangmeng
Zhang, Shichao
contents Deep neural networks enriched with structural information have been widely employed for facial expression recognition tasks. However, these methods often depend on hierarchical information rather than face property to finish expression recognition. In this paper, we propose a cross-modal network with strong biological and structural information for facial expression recognition (CMNet). CMNet can respectively learn expression information via face symmetry on a whole face, left and right half faces to extract complementary facial features. To prevent negative effect of biological and structural information fusion, a salient facial information refinement module can obtain salient facial expression information to improve stability of an obtained facial expression classifier. To reduce reliance on unilateral facial features, a half-face alignment optimization mechanism is designed to align obtained expression information of learned left and right half faces. Our experimental results demonstrate that CMNet outperforms several novel methods, i.e., SCN and LAENet-SA for facial expression recognition. Codes can be obtained at https://github.com/hellloxiaotian/CMNet.
format Preprint
id arxiv_https___arxiv_org_abs_2605_04439
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle A cross-modal network for facial expression recognition
Tian, Chunwei
Xie, Jingyuan
Zhang, Qi
Li, Chao
Zuo, Wangmeng
Zhang, Shichao
Computer Vision and Pattern Recognition
Deep neural networks enriched with structural information have been widely employed for facial expression recognition tasks. However, these methods often depend on hierarchical information rather than face property to finish expression recognition. In this paper, we propose a cross-modal network with strong biological and structural information for facial expression recognition (CMNet). CMNet can respectively learn expression information via face symmetry on a whole face, left and right half faces to extract complementary facial features. To prevent negative effect of biological and structural information fusion, a salient facial information refinement module can obtain salient facial expression information to improve stability of an obtained facial expression classifier. To reduce reliance on unilateral facial features, a half-face alignment optimization mechanism is designed to align obtained expression information of learned left and right half faces. Our experimental results demonstrate that CMNet outperforms several novel methods, i.e., SCN and LAENet-SA for facial expression recognition. Codes can be obtained at https://github.com/hellloxiaotian/CMNet.
title A cross-modal network for facial expression recognition
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.04439