EgoReAct: Egocentric Video-Driven 3D Human Reaction Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Libo, Li, Zekun, Li, Tianyu, Cao, Zeyu, Xu, Rui, Long, Xiaoxiao, Wang, Wenjia, Wang, Jingbo, Liu, Yuan, Wang, Wenping, Zhou, Daquan, Komura, Taku, Dou, Zhiyang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912801403961344
author Zhang, Libo
Li, Zekun
Li, Tianyu
Cao, Zeyu
Xu, Rui
Long, Xiaoxiao
Wang, Wenjia
Wang, Jingbo
Liu, Yuan
Wang, Wenping
Zhou, Daquan
Komura, Taku
Dou, Zhiyang
author_facet Zhang, Libo
Li, Zekun
Li, Tianyu
Cao, Zeyu
Xu, Rui
Long, Xiaoxiao
Wang, Wenjia
Wang, Jingbo
Liu, Yuan
Wang, Wenping
Zhou, Daquan
Komura, Taku
Dou, Zhiyang
contents Humans exhibit adaptive, context-sensitive responses to egocentric visual input. However, faithfully modeling such reactions from egocentric video remains challenging due to the dual requirements of strictly causal generation and precise 3D spatial alignment. To tackle this problem, we first construct the Human Reaction Dataset (HRD) to address data scarcity and misalignment by building a spatially aligned egocentric video-reaction dataset, as existing datasets (e.g., ViMo) suffer from significant spatial inconsistency between the egocentric video and reaction motion, e.g., dynamically moving motions are always paired with fixed-camera videos. Leveraging HRD, we present EgoReAct, the first autoregressive framework that generates 3D-aligned human reaction motions from egocentric video streams in real-time. We first compress the reaction motion into a compact yet expressive latent space via a Vector Quantised-Variational AutoEncoder and then train a Generative Pre-trained Transformer for reaction generation from the visual input. EgoReAct incorporates 3D dynamic features, i.e., metric depth, and head dynamics during the generation, which effectively enhance spatial grounding. Extensive experiments demonstrate that EgoReAct achieves remarkably higher realism, spatial consistency, and generation efficiency compared with prior methods, while maintaining strict causality during generation. We will release code, models, and data upon acceptance.
format Preprint
id arxiv_https___arxiv_org_abs_2512_22808
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EgoReAct: Egocentric Video-Driven 3D Human Reaction Generation
Zhang, Libo
Li, Zekun
Li, Tianyu
Cao, Zeyu
Xu, Rui
Long, Xiaoxiao
Wang, Wenjia
Wang, Jingbo
Liu, Yuan
Wang, Wenping
Zhou, Daquan
Komura, Taku
Dou, Zhiyang
Computer Vision and Pattern Recognition
Artificial Intelligence
Humans exhibit adaptive, context-sensitive responses to egocentric visual input. However, faithfully modeling such reactions from egocentric video remains challenging due to the dual requirements of strictly causal generation and precise 3D spatial alignment. To tackle this problem, we first construct the Human Reaction Dataset (HRD) to address data scarcity and misalignment by building a spatially aligned egocentric video-reaction dataset, as existing datasets (e.g., ViMo) suffer from significant spatial inconsistency between the egocentric video and reaction motion, e.g., dynamically moving motions are always paired with fixed-camera videos. Leveraging HRD, we present EgoReAct, the first autoregressive framework that generates 3D-aligned human reaction motions from egocentric video streams in real-time. We first compress the reaction motion into a compact yet expressive latent space via a Vector Quantised-Variational AutoEncoder and then train a Generative Pre-trained Transformer for reaction generation from the visual input. EgoReAct incorporates 3D dynamic features, i.e., metric depth, and head dynamics during the generation, which effectively enhance spatial grounding. Extensive experiments demonstrate that EgoReAct achieves remarkably higher realism, spatial consistency, and generation efficiency compared with prior methods, while maintaining strict causality during generation. We will release code, models, and data upon acceptance.
title EgoReAct: Egocentric Video-Driven 3D Human Reaction Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2512.22808