HERO: Human Reaction Generation from Videos

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, Chengjun, Zhai, Wei, Yang, Yuhang, Cao, Yang, Zha, Zheng-Jun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916649675784192
author Yu, Chengjun
Zhai, Wei
Yang, Yuhang
Cao, Yang
Zha, Zheng-Jun
author_facet Yu, Chengjun
Zhai, Wei
Yang, Yuhang
Cao, Yang
Zha, Zheng-Jun
contents Human reaction generation represents a significant research domain for interactive AI, as humans constantly interact with their surroundings. Previous works focus mainly on synthesizing the reactive motion given a human motion sequence. This paradigm limits interaction categories to human-human interactions and ignores emotions that may influence reaction generation. In this work, we propose to generate 3D human reactions from RGB videos, which involves a wider range of interaction categories and naturally provides information about expressions that may reflect the subject's emotions. To cope with this task, we present HERO, a simple yet powerful framework for Human rEaction geneRation from videOs. HERO considers both global and frame-level local representations of the video to extract the interaction intention, and then uses the extracted interaction intention to guide the synthesis of the reaction. Besides, local visual representations are continuously injected into the model to maximize the exploitation of the dynamic properties inherent in videos. Furthermore, the ViMo dataset containing paired Video-Motion data is collected to support the task. In addition to human-human interactions, these video-motion pairs also cover animal-human interactions and scene-human interactions. Extensive experiments demonstrate the superiority of our methodology. The code and dataset will be publicly available at https://jackyu6.github.io/HERO.
format Preprint
id arxiv_https___arxiv_org_abs_2503_08270
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HERO: Human Reaction Generation from Videos
Yu, Chengjun
Zhai, Wei
Yang, Yuhang
Cao, Yang
Zha, Zheng-Jun
Computer Vision and Pattern Recognition
Human reaction generation represents a significant research domain for interactive AI, as humans constantly interact with their surroundings. Previous works focus mainly on synthesizing the reactive motion given a human motion sequence. This paradigm limits interaction categories to human-human interactions and ignores emotions that may influence reaction generation. In this work, we propose to generate 3D human reactions from RGB videos, which involves a wider range of interaction categories and naturally provides information about expressions that may reflect the subject's emotions. To cope with this task, we present HERO, a simple yet powerful framework for Human rEaction geneRation from videOs. HERO considers both global and frame-level local representations of the video to extract the interaction intention, and then uses the extracted interaction intention to guide the synthesis of the reaction. Besides, local visual representations are continuously injected into the model to maximize the exploitation of the dynamic properties inherent in videos. Furthermore, the ViMo dataset containing paired Video-Motion data is collected to support the task. In addition to human-human interactions, these video-motion pairs also cover animal-human interactions and scene-human interactions. Extensive experiments demonstrate the superiority of our methodology. The code and dataset will be publicly available at https://jackyu6.github.io/HERO.
title HERO: Human Reaction Generation from Videos
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.08270