SARAH: Spatially Aware Real-time Agentic Humans

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ng, Evonne, Zhang, Siwei, Chen, Zhang, Zollhoefer, Michael, Richard, Alexander
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914340117938176
author Ng, Evonne
Zhang, Siwei
Chen, Zhang
Zollhoefer, Michael
Richard, Alexander
author_facet Ng, Evonne
Zhang, Siwei
Chen, Zhang
Zollhoefer, Michael
Richard, Alexander
contents As embodied agents become central to VR, telepresence, and digital human applications, their motion must go beyond speech-aligned gestures: agents should turn toward users, respond to their movement, and maintain natural gaze. Current methods lack this spatial awareness. We close this gap with the first real-time, fully causal method for spatially-aware conversational motion, deployable on a streaming VR headset. Given a user's position and dyadic audio, our approach produces full-body motion that aligns gestures with speech while orienting the agent according to the user. Our architecture combines a causal transformer-based VAE with interleaved latent tokens for streaming inference and a flow matching model conditioned on user trajectory and audio. To support varying gaze preferences, we introduce a gaze scoring mechanism with classifier-free guidance to decouple learning from control: the model captures natural spatial alignment from data, while users can adjust eye contact intensity at inference time. On the Embody 3D dataset, our method achieves state-of-the-art motion quality at over 300 FPS -- 3x faster than non-causal baselines -- while capturing the subtle spatial dynamics of natural conversation. We validate our approach on a live VR system, bringing spatially-aware conversational agents to real-time deployment. Please see https://evonneng.github.io/sarah/ for details.
format Preprint
id arxiv_https___arxiv_org_abs_2602_18432
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SARAH: Spatially Aware Real-time Agentic Humans
Ng, Evonne
Zhang, Siwei
Chen, Zhang
Zollhoefer, Michael
Richard, Alexander
Computer Vision and Pattern Recognition
As embodied agents become central to VR, telepresence, and digital human applications, their motion must go beyond speech-aligned gestures: agents should turn toward users, respond to their movement, and maintain natural gaze. Current methods lack this spatial awareness. We close this gap with the first real-time, fully causal method for spatially-aware conversational motion, deployable on a streaming VR headset. Given a user's position and dyadic audio, our approach produces full-body motion that aligns gestures with speech while orienting the agent according to the user. Our architecture combines a causal transformer-based VAE with interleaved latent tokens for streaming inference and a flow matching model conditioned on user trajectory and audio. To support varying gaze preferences, we introduce a gaze scoring mechanism with classifier-free guidance to decouple learning from control: the model captures natural spatial alignment from data, while users can adjust eye contact intensity at inference time. On the Embody 3D dataset, our method achieves state-of-the-art motion quality at over 300 FPS -- 3x faster than non-causal baselines -- while capturing the subtle spatial dynamics of natural conversation. We validate our approach on a live VR system, bringing spatially-aware conversational agents to real-time deployment. Please see https://evonneng.github.io/sarah/ for details.
title SARAH: Spatially Aware Real-time Agentic Humans
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.18432