HANDI: Hand-Centric Text-and-Image Conditioned Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Yayuan, Cao, Zhi, Corso, Jason J.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909687553720320
author Li, Yayuan
Cao, Zhi
Corso, Jason J.
author_facet Li, Yayuan
Cao, Zhi
Corso, Jason J.
contents Despite the recent strides in video generation, state-of-the-art methods still struggle with elements of visual detail. One particularly challenging case is the class of videos in which the intricate motion of the hand coupled with a mostly stable and otherwise distracting environment is necessary to convey the execution of some complex action and its effects. To address these challenges, we introduce a new method for video generation that focuses on hand-centric actions. Our diffusion-based method incorporates two distinct innovations. First, we propose an automatic method to generate the motion area -- the region in the video in which the detailed activities occur -- guided by both the visual context and the action text prompt, rather than assuming this region can be provided manually as is now commonplace. Second, we introduce a critical Hand Refinement Loss to guide the diffusion model to focus on smooth and consistent hand poses. We evaluate our method on challenging augmented datasets based on EpicKitchens and Ego4D, demonstrating significant improvements over state-of-the-art methods in terms of action clarity, especially of the hand motion in the target region, across diverse environments and actions. Video results can be found in https://excitedbutter.github.io/project_page
format Preprint
id arxiv_https___arxiv_org_abs_2412_04189
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle HANDI: Hand-Centric Text-and-Image Conditioned Video Generation
Li, Yayuan
Cao, Zhi
Corso, Jason J.
Computer Vision and Pattern Recognition
Despite the recent strides in video generation, state-of-the-art methods still struggle with elements of visual detail. One particularly challenging case is the class of videos in which the intricate motion of the hand coupled with a mostly stable and otherwise distracting environment is necessary to convey the execution of some complex action and its effects. To address these challenges, we introduce a new method for video generation that focuses on hand-centric actions. Our diffusion-based method incorporates two distinct innovations. First, we propose an automatic method to generate the motion area -- the region in the video in which the detailed activities occur -- guided by both the visual context and the action text prompt, rather than assuming this region can be provided manually as is now commonplace. Second, we introduce a critical Hand Refinement Loss to guide the diffusion model to focus on smooth and consistent hand poses. We evaluate our method on challenging augmented datasets based on EpicKitchens and Ego4D, demonstrating significant improvements over state-of-the-art methods in terms of action clarity, especially of the hand motion in the target region, across diverse environments and actions. Video results can be found in https://excitedbutter.github.io/project_page
title HANDI: Hand-Centric Text-and-Image Conditioned Video Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.04189