Improving Personalized Image Generation through Social Context Feedback

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gupta, Parul, Dhall, Abhinav, Do, Thanh-Toan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909698501902336
author Gupta, Parul
Dhall, Abhinav
Do, Thanh-Toan
author_facet Gupta, Parul
Dhall, Abhinav
Do, Thanh-Toan
contents Personalized image generation, where reference images of one or more subjects are used to generate their image according to a scene description, has gathered significant interest in the community. However, such generated images suffer from three major limitations -- complex activities, such as $<$man, pushing, motorcycle$>$ are not generated properly with incorrect human poses, reference human identities are not preserved, and generated human gaze patterns are unnatural/inconsistent with the scene description. In this work, we propose to overcome these shortcomings through feedback-based fine-tuning of existing personalized generation methods, wherein, state-of-art detectors of pose, human-object-interaction, human facial recognition and human gaze-point estimation are used to refine the diffusion model. We also propose timestep-based inculcation of different feedback modules, depending upon whether the signal is low-level (such as human pose), or high-level (such as gaze point). The images generated in this manner show an improvement in the generated interactions, facial identities and image quality over three benchmark datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2507_16095
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Improving Personalized Image Generation through Social Context Feedback
Gupta, Parul
Dhall, Abhinav
Do, Thanh-Toan
Computer Vision and Pattern Recognition
Personalized image generation, where reference images of one or more subjects are used to generate their image according to a scene description, has gathered significant interest in the community. However, such generated images suffer from three major limitations -- complex activities, such as $<$man, pushing, motorcycle$>$ are not generated properly with incorrect human poses, reference human identities are not preserved, and generated human gaze patterns are unnatural/inconsistent with the scene description. In this work, we propose to overcome these shortcomings through feedback-based fine-tuning of existing personalized generation methods, wherein, state-of-art detectors of pose, human-object-interaction, human facial recognition and human gaze-point estimation are used to refine the diffusion model. We also propose timestep-based inculcation of different feedback modules, depending upon whether the signal is low-level (such as human pose), or high-level (such as gaze point). The images generated in this manner show an improvement in the generated interactions, facial identities and image quality over three benchmark datasets.
title Improving Personalized Image Generation through Social Context Feedback
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.16095