VerbDiff: Text-Only Diffusion Models with Enhanced Interaction Awareness

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cha, SeungJu, Lee, Kwanyoung, Kim, Ye-Chan, Oh, Hyunwoo, Kim, Dong-Jin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908276180910080
author Cha, SeungJu
Lee, Kwanyoung
Kim, Ye-Chan
Oh, Hyunwoo
Kim, Dong-Jin
author_facet Cha, SeungJu
Lee, Kwanyoung
Kim, Ye-Chan
Oh, Hyunwoo
Kim, Dong-Jin
contents Recent large-scale text-to-image diffusion models generate photorealistic images but often struggle to accurately depict interactions between humans and objects due to their limited ability to differentiate various interaction words. In this work, we propose VerbDiff to address the challenge of capturing nuanced interactions within text-to-image diffusion models. VerbDiff is a novel text-to-image generation model that weakens the bias between interaction words and objects, enhancing the understanding of interactions. Specifically, we disentangle various interaction words from frequency-based anchor words and leverage localized interaction regions from generated images to help the model better capture semantics in distinctive words without extra conditions. Our approach enables the model to accurately understand the intended interaction between humans and objects, producing high-quality images with accurate interactions aligned with specified verbs. Extensive experiments on the HICO-DET dataset demonstrate the effectiveness of our method compared to previous approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2503_16406
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VerbDiff: Text-Only Diffusion Models with Enhanced Interaction Awareness
Cha, SeungJu
Lee, Kwanyoung
Kim, Ye-Chan
Oh, Hyunwoo
Kim, Dong-Jin
Graphics
Computer Vision and Pattern Recognition
Multimedia
Recent large-scale text-to-image diffusion models generate photorealistic images but often struggle to accurately depict interactions between humans and objects due to their limited ability to differentiate various interaction words. In this work, we propose VerbDiff to address the challenge of capturing nuanced interactions within text-to-image diffusion models. VerbDiff is a novel text-to-image generation model that weakens the bias between interaction words and objects, enhancing the understanding of interactions. Specifically, we disentangle various interaction words from frequency-based anchor words and leverage localized interaction regions from generated images to help the model better capture semantics in distinctive words without extra conditions. Our approach enables the model to accurately understand the intended interaction between humans and objects, producing high-quality images with accurate interactions aligned with specified verbs. Extensive experiments on the HICO-DET dataset demonstrate the effectiveness of our method compared to previous approaches.
title VerbDiff: Text-Only Diffusion Models with Enhanced Interaction Awareness
topic Graphics
Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2503.16406