ARMFlow: AutoRegressive MeanFlow for Online 3D Human Reaction Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Geng, Zichen, Hayder, Zeeshan, Liu, Wei, Wang, Hesheng, Mian, Ajmal
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910052284104704
author Geng, Zichen
Hayder, Zeeshan
Liu, Wei
Wang, Hesheng
Mian, Ajmal
author_facet Geng, Zichen
Hayder, Zeeshan
Liu, Wei
Wang, Hesheng
Mian, Ajmal
contents 3D human reaction generation faces three main challenges:(1) high motion fidelity, (2) real-time inference, and (3) autoregressive adaptability for online scenarios. Existing methods fail to meet all three simultaneously. We propose ARMFlow, a MeanFlow-based autoregressive framework that models temporal dependencies between actor and reactor motions. It consists of a causal context encoder and an MLP-based velocity predictor. We introduce Bootstrap Contextual Encoding (BSCE) in training, encoding generated history instead of the ground-truth ones, to alleviate error accumulation in autoregressive generation. We further introduce the offline variant ReMFlow, achieving state-of-the-art performance with the fastest inference among offline methods. Our ARMFlow addresses key limitations of online settings by: (1) enhancing semantic alignment via a global contextual encoder; (2) achieving high accuracy and low latency in a single-step inference; and (3) reducing accumulated errors through BSCE. Our single-step online generation surpasses existing online methods on InterHuman and InterX by about 30% in FID, while matching offline state-of-the-art performance despite using only partial sequence conditions.
format Preprint
id arxiv_https___arxiv_org_abs_2512_16234
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ARMFlow: AutoRegressive MeanFlow for Online 3D Human Reaction Generation
Geng, Zichen
Hayder, Zeeshan
Liu, Wei
Wang, Hesheng
Mian, Ajmal
Computer Vision and Pattern Recognition
3D human reaction generation faces three main challenges:(1) high motion fidelity, (2) real-time inference, and (3) autoregressive adaptability for online scenarios. Existing methods fail to meet all three simultaneously. We propose ARMFlow, a MeanFlow-based autoregressive framework that models temporal dependencies between actor and reactor motions. It consists of a causal context encoder and an MLP-based velocity predictor. We introduce Bootstrap Contextual Encoding (BSCE) in training, encoding generated history instead of the ground-truth ones, to alleviate error accumulation in autoregressive generation. We further introduce the offline variant ReMFlow, achieving state-of-the-art performance with the fastest inference among offline methods. Our ARMFlow addresses key limitations of online settings by: (1) enhancing semantic alignment via a global contextual encoder; (2) achieving high accuracy and low latency in a single-step inference; and (3) reducing accumulated errors through BSCE. Our single-step online generation surpasses existing online methods on InterHuman and InterX by about 30% in FID, while matching offline state-of-the-art performance despite using only partial sequence conditions.
title ARMFlow: AutoRegressive MeanFlow for Online 3D Human Reaction Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.16234