ARGen: Affect-Reinforced Generative Augmentation towards Vision-based Dynamic Emotion Perception

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Huanzhen, Zhou, Ziheng, Song, Jiaqi, He, Li, Lan, Yunshi, Wang, Yan, Zhang, Wenqiang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913029287837696
author Wang, Huanzhen
Zhou, Ziheng
Song, Jiaqi
He, Li
Lan, Yunshi
Wang, Yan
Zhang, Wenqiang
author_facet Wang, Huanzhen
Zhou, Ziheng
Song, Jiaqi
He, Li
Lan, Yunshi
Wang, Yan
Zhang, Wenqiang
contents Dynamic facial expression recognition in the wild remains challenging due to data scarcity and long-tail distributions, which hinder models from effectively learning the temporal dynamics of scarce emotions. To address these limitations, we propose ARGen, an Affect-Reinforced Generative Augmentation Framework that enables data-adaptive dynamic expression generation for robust emotion perception. ARGen operates in two stages: Affective Semantic Injection (ASI) and Adaptive Reinforcement Diffusion (ARD). The ASI stage establishes affective knowledge alignment through facial Action Units and employs a retrieval-augmented prompt generation strategy to synthesize consistent and fine-grained affective descriptions via large-scale visual-language models, thereby injecting interpretable emotional priors into the generation process. The ARD stage integrates text-conditioned image-to-video diffusion with reinforcement learning, introducing inter-frame conditional guidance and a multi-objective reward function to jointly optimize expression naturalness, facial integrity, and generative efficiency. Extensive experiments on both generation and recognition tasks verify that ARGen substantially enhances synthesis fidelity and improves recognition performance, establishing an interpretable and generalizable generative augmentation paradigm for vision-based affective computing.
format Preprint
id arxiv_https___arxiv_org_abs_2604_12255
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ARGen: Affect-Reinforced Generative Augmentation towards Vision-based Dynamic Emotion Perception
Wang, Huanzhen
Zhou, Ziheng
Song, Jiaqi
He, Li
Lan, Yunshi
Wang, Yan
Zhang, Wenqiang
Computer Vision and Pattern Recognition
Artificial Intelligence
Dynamic facial expression recognition in the wild remains challenging due to data scarcity and long-tail distributions, which hinder models from effectively learning the temporal dynamics of scarce emotions. To address these limitations, we propose ARGen, an Affect-Reinforced Generative Augmentation Framework that enables data-adaptive dynamic expression generation for robust emotion perception. ARGen operates in two stages: Affective Semantic Injection (ASI) and Adaptive Reinforcement Diffusion (ARD). The ASI stage establishes affective knowledge alignment through facial Action Units and employs a retrieval-augmented prompt generation strategy to synthesize consistent and fine-grained affective descriptions via large-scale visual-language models, thereby injecting interpretable emotional priors into the generation process. The ARD stage integrates text-conditioned image-to-video diffusion with reinforcement learning, introducing inter-frame conditional guidance and a multi-objective reward function to jointly optimize expression naturalness, facial integrity, and generative efficiency. Extensive experiments on both generation and recognition tasks verify that ARGen substantially enhances synthesis fidelity and improves recognition performance, establishing an interpretable and generalizable generative augmentation paradigm for vision-based affective computing.
title ARGen: Affect-Reinforced Generative Augmentation towards Vision-based Dynamic Emotion Perception
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2604.12255