Emotion-Director: Bridging Affective Shortcut in Emotion-Oriented Image Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jia, Guoli, Hu, Junyao, Long, Xinwei, Tian, Kai, Zhang, Kaiyan, Zhao, KaiKai, Ding, Ning, Zhou, Bowen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911332596449280
author Jia, Guoli
Hu, Junyao
Long, Xinwei
Tian, Kai
Zhang, Kaiyan
Zhao, KaiKai
Ding, Ning
Zhou, Bowen
author_facet Jia, Guoli
Hu, Junyao
Long, Xinwei
Tian, Kai
Zhang, Kaiyan
Zhao, KaiKai
Ding, Ning
Zhou, Bowen
contents Image generation based on diffusion models has demonstrated impressive capability, motivating exploration into diverse and specialized applications. Owing to the importance of emotion in advertising, emotion-oriented image generation has attracted increasing attention. However, current emotion-oriented methods suffer from an affective shortcut, where emotions are approximated to semantics. As evidenced by two decades of research, emotion is not equivalent to semantics. To this end, we propose Emotion-Director, a cross-modal collaboration framework consisting of two modules. First, we propose a cross-Modal Collaborative diffusion model, abbreviated as MC-Diffusion. MC-Diffusion integrates visual prompts with textual prompts for guidance, enabling the generation of emotion-oriented images beyond semantics. Further, we improve the DPO optimization by a negative visual prompt, enhancing the model's sensitivity to different emotions under the same semantics. Second, we propose MC-Agent, a cross-Modal Collaborative Agent system that rewrites textual prompts to express the intended emotions. To avoid template-like rewrites, MC-Agent employs multi-agents to simulate human subjectivity toward emotions, and adopts a chain-of-concept workflow that improves the visual expressiveness of the rewritten prompts. Extensive qualitative and quantitative experiments demonstrate the superiority of Emotion-Director in emotion-oriented image generation.
format Preprint
id arxiv_https___arxiv_org_abs_2512_19479
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Emotion-Director: Bridging Affective Shortcut in Emotion-Oriented Image Generation
Jia, Guoli
Hu, Junyao
Long, Xinwei
Tian, Kai
Zhang, Kaiyan
Zhao, KaiKai
Ding, Ning
Zhou, Bowen
Computer Vision and Pattern Recognition
Image generation based on diffusion models has demonstrated impressive capability, motivating exploration into diverse and specialized applications. Owing to the importance of emotion in advertising, emotion-oriented image generation has attracted increasing attention. However, current emotion-oriented methods suffer from an affective shortcut, where emotions are approximated to semantics. As evidenced by two decades of research, emotion is not equivalent to semantics. To this end, we propose Emotion-Director, a cross-modal collaboration framework consisting of two modules. First, we propose a cross-Modal Collaborative diffusion model, abbreviated as MC-Diffusion. MC-Diffusion integrates visual prompts with textual prompts for guidance, enabling the generation of emotion-oriented images beyond semantics. Further, we improve the DPO optimization by a negative visual prompt, enhancing the model's sensitivity to different emotions under the same semantics. Second, we propose MC-Agent, a cross-Modal Collaborative Agent system that rewrites textual prompts to express the intended emotions. To avoid template-like rewrites, MC-Agent employs multi-agents to simulate human subjectivity toward emotions, and adopts a chain-of-concept workflow that improves the visual expressiveness of the rewritten prompts. Extensive qualitative and quantitative experiments demonstrate the superiority of Emotion-Director in emotion-oriented image generation.
title Emotion-Director: Bridging Affective Shortcut in Emotion-Oriented Image Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.19479