Solution to the 10th ABAW Expression Recognition Challenge: A Robust Multimodal Framework with Safe Cross-Attention and Modality Dropout

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yu, Jun, Zheng, Naixiang, Wang, Guoyuan, Zhang, Yunxiang, Zhu, Lingsi, Liang, Jiaen, Huang, Wei, Liu, Shengping
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912954696335360
author Yu, Jun
Zheng, Naixiang
Wang, Guoyuan
Zhang, Yunxiang
Zhu, Lingsi
Liang, Jiaen
Huang, Wei
Liu, Shengping
author_facet Yu, Jun
Zheng, Naixiang
Wang, Guoyuan
Zhang, Yunxiang
Zhu, Lingsi
Liang, Jiaen
Huang, Wei
Liu, Shengping
contents Emotion recognition in real-world environments is hindered by partial occlusions, missing modalities, and severe class imbalance. To address these issues, particularly for the Affective Behavior Analysis in-the-wild (ABAW) Expression challenge, we propose a multimodal framework that dynamically fuses visual and audio representations. Our approach uses a dual-branch Transformer architecture featuring a safe cross-attention mechanism and a modality dropout strategy. This design allows the network to rely on audio-based predictions when visual cues are absent. To mitigate the long-tail distribution of the Aff-Wild2 dataset, we apply focal loss optimization, combined with a sliding-window soft voting strategy to capture dynamic emotional transitions and reduce frame-level classification jitter. Experiments demonstrate that our framework effectively handles missing modalities and complex spatiotemporal dependencies, achieving an accuracy of 60.79% and an F1-score of 0.5029 on the Aff-Wild2 validation set.
format Preprint
id arxiv_https___arxiv_org_abs_2603_08034
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Solution to the 10th ABAW Expression Recognition Challenge: A Robust Multimodal Framework with Safe Cross-Attention and Modality Dropout
Yu, Jun
Zheng, Naixiang
Wang, Guoyuan
Zhang, Yunxiang
Zhu, Lingsi
Liang, Jiaen
Huang, Wei
Liu, Shengping
Computer Vision and Pattern Recognition
Artificial Intelligence
Emotion recognition in real-world environments is hindered by partial occlusions, missing modalities, and severe class imbalance. To address these issues, particularly for the Affective Behavior Analysis in-the-wild (ABAW) Expression challenge, we propose a multimodal framework that dynamically fuses visual and audio representations. Our approach uses a dual-branch Transformer architecture featuring a safe cross-attention mechanism and a modality dropout strategy. This design allows the network to rely on audio-based predictions when visual cues are absent. To mitigate the long-tail distribution of the Aff-Wild2 dataset, we apply focal loss optimization, combined with a sliding-window soft voting strategy to capture dynamic emotional transitions and reduce frame-level classification jitter. Experiments demonstrate that our framework effectively handles missing modalities and complex spatiotemporal dependencies, achieving an accuracy of 60.79% and an F1-score of 0.5029 on the Aff-Wild2 validation set.
title Solution to the 10th ABAW Expression Recognition Challenge: A Robust Multimodal Framework with Safe Cross-Attention and Modality Dropout
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2603.08034