Learning Complex Non-Rigid Image Edits from Multimodal Conditioning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Warner, Nikolai, Kolb, Jack, Hahn, Meera, Birodkar, Vighnesh, Huang, Jonathan, Essa, Irfan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929628943220736
author Warner, Nikolai
Kolb, Jack
Hahn, Meera
Birodkar, Vighnesh
Huang, Jonathan
Essa, Irfan
author_facet Warner, Nikolai
Kolb, Jack
Hahn, Meera
Birodkar, Vighnesh
Huang, Jonathan
Essa, Irfan
contents In this paper we focus on inserting a given human (specifically, a single image of a person) into a novel scene. Our method, which builds on top of Stable Diffusion, yields natural looking images while being highly controllable with text and pose. To accomplish this we need to train on pairs of images, the first a reference image with the person, the second a "target image" showing the same person (with a different pose and possibly in a different background). Additionally we require a text caption describing the new pose relative to that in the reference image. In this paper we present a novel dataset following this criteria, which we create using pairs of frames from human-centric and action-rich videos and employing a multimodal LLM to automatically summarize the difference in human pose for the text captions. We demonstrate that identity preservation is a more challenging task in scenes "in-the-wild", and especially scenes where there is an interaction between persons and objects. Combining the weak supervision from noisy captions, with robust 2D pose improves the quality of person-object interactions.
format Preprint
id arxiv_https___arxiv_org_abs_2412_10219
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Learning Complex Non-Rigid Image Edits from Multimodal Conditioning
Warner, Nikolai
Kolb, Jack
Hahn, Meera
Birodkar, Vighnesh
Huang, Jonathan
Essa, Irfan
Computer Vision and Pattern Recognition
In this paper we focus on inserting a given human (specifically, a single image of a person) into a novel scene. Our method, which builds on top of Stable Diffusion, yields natural looking images while being highly controllable with text and pose. To accomplish this we need to train on pairs of images, the first a reference image with the person, the second a "target image" showing the same person (with a different pose and possibly in a different background). Additionally we require a text caption describing the new pose relative to that in the reference image. In this paper we present a novel dataset following this criteria, which we create using pairs of frames from human-centric and action-rich videos and employing a multimodal LLM to automatically summarize the difference in human pose for the text captions. We demonstrate that identity preservation is a more challenging task in scenes "in-the-wild", and especially scenes where there is an interaction between persons and objects. Combining the weak supervision from noisy captions, with robust 2D pose improves the quality of person-object interactions.
title Learning Complex Non-Rigid Image Edits from Multimodal Conditioning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.10219