SARAH: Spatially Aware Real-time Agentic Humans
Fuente:
arXiv
Saved in:
| Main Authors: | Ng, Evonne, Zhang, Siwei, Chen, Zhang, Zollhoefer, Michael, Richard, Alexander |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AV-Flow: Transforming Text to Audio-Visual Human-like Interactions
by: Chatziagapi, Aggelina, et al.
Published: (2025)
by: Chatziagapi, Aggelina, et al.
Published: (2025)
From Audio to Photoreal Embodiment: Synthesizing Humans in Conversations
by: Ng, Evonne, et al.
Published: (2024)
by: Ng, Evonne, et al.
Published: (2024)
DuoMo: Dual Motion Diffusion for World-Space Human Reconstruction
by: Wang, Yufu, et al.
Published: (2026)
by: Wang, Yufu, et al.
Published: (2026)
Embody 3D: A Large-scale Multimodal Motion and Behavior Dataset
by: McLean, Claire, et al.
Published: (2025)
by: McLean, Claire, et al.
Published: (2025)
Pose Priors from Language Models
by: Subramanian, Sanjay, et al.
Published: (2024)
by: Subramanian, Sanjay, et al.
Published: (2024)
3D Human Pose Estimation via Spatial Graph Order Attention and Temporal Body Aware Transformer
by: Aouaidjia, Kamel, et al.
Published: (2025)
by: Aouaidjia, Kamel, et al.
Published: (2025)
Diffusion Forcing for Multi-Agent Interaction Sequence Modeling
by: Maluleke, Vongani H., et al.
Published: (2025)
by: Maluleke, Vongani H., et al.
Published: (2025)
RoHM: Robust Human Motion Reconstruction via Diffusion
by: Zhang, Siwei, et al.
Published: (2024)
by: Zhang, Siwei, et al.
Published: (2024)
Spatial-Aware Self-Supervision for Medical 3D Imaging with Multi-Granularity Observable Tasks
by: Zhang, Yiqin, et al.
Published: (2025)
by: Zhang, Yiqin, et al.
Published: (2025)
End-to-End Spatial-Temporal Transformer for Real-time 4D HOI Reconstruction
by: Zhang, Haoyu, et al.
Published: (2026)
by: Zhang, Haoyu, et al.
Published: (2026)
A Semantic-Aware and Multi-Guided Network for Infrared-Visible Image Fusion
by: Zhang, Xiaoli, et al.
Published: (2024)
by: Zhang, Xiaoli, et al.
Published: (2024)
Masked Modeling for Human Motion Recovery Under Occlusions
by: Qian, Zhiyin, et al.
Published: (2026)
by: Qian, Zhiyin, et al.
Published: (2026)
VTAgent: Agentic Keyframe Anchoring for Evidence-Aware Video TextVQA
by: He, Haibin, et al.
Published: (2026)
by: He, Haibin, et al.
Published: (2026)
PhyMAGIC: Physical Motion-Aware Generative Inference with Confidence-guided LLM
by: Meng, Siwei, et al.
Published: (2025)
by: Meng, Siwei, et al.
Published: (2025)
Human Gaussian Splatting: Real-time Rendering of Animatable Avatars
by: Moreau, Arthur, et al.
Published: (2023)
by: Moreau, Arthur, et al.
Published: (2023)
Patho-AgenticRAG: Towards Multimodal Agentic Retrieval-Augmented Generation for Pathology VLMs via Reinforcement Learning
by: Zhang, Wenchuan, et al.
Published: (2025)
by: Zhang, Wenchuan, et al.
Published: (2025)
Active Exploring like a Pigeon: Reinforcing Spatial Reasoning via Agentic Vision-Language Models
by: Deng, Wei, et al.
Published: (2026)
by: Deng, Wei, et al.
Published: (2026)
Real-time High-fidelity Gaussian Human Avatars with Position-based Interpolation of Spatially Distributed MLPs
by: Zhan, Youyi, et al.
Published: (2025)
by: Zhan, Youyi, et al.
Published: (2025)
Egocentric Visibility-Aware Human Pose Estimation
by: Dai, Peng, et al.
Published: (2026)
by: Dai, Peng, et al.
Published: (2026)
TACR-YOLO: A Real-time Detection Framework for Abnormal Human Behaviors Enhanced with Coordinate and Task-Aware Representations
by: Yin, Xinyi, et al.
Published: (2025)
by: Yin, Xinyi, et al.
Published: (2025)
Real-time Spatial-temporal Traversability Assessment via Feature-based Sparse Gaussian Process
by: Hou, Zhenyu, et al.
Published: (2025)
by: Hou, Zhenyu, et al.
Published: (2025)
Fusing Pixels and Genes: Spatially-Aware Learning in Computational Pathology
by: Han, Minghao, et al.
Published: (2026)
by: Han, Minghao, et al.
Published: (2026)
LGM-Pose: A Lightweight Global Modeling Network for Real-time Human Pose Estimation
by: Guo, Biao, et al.
Published: (2025)
by: Guo, Biao, et al.
Published: (2025)
SPC-NeRF: Spatial Predictive Compression for Voxel Based Radiance Field
by: Song, Zetian, et al.
Published: (2024)
by: Song, Zetian, et al.
Published: (2024)
OARS: Process-Aware Online Alignment for Generative Real-World Image Super-Resolution
by: Zhao, Shijie, et al.
Published: (2026)
by: Zhao, Shijie, et al.
Published: (2026)
EditRefiner: A Human-Aligned Agentic Framework for Image Editing Refinement
by: Xu, Zitong, et al.
Published: (2026)
by: Xu, Zitong, et al.
Published: (2026)
GaussianGAN: Real-Time Photorealistic controllable Human Avatars
by: Lakhal, Mohamed Ilyes, et al.
Published: (2025)
by: Lakhal, Mohamed Ilyes, et al.
Published: (2025)
PARSE: Part-Aware Relational Spatial Modeling
by: Bai, Yinuo, et al.
Published: (2026)
by: Bai, Yinuo, et al.
Published: (2026)
MoSA: Motion-Coherent Human Video Generation via Structure-Appearance Decoupling
by: Wang, Haoyu, et al.
Published: (2025)
by: Wang, Haoyu, et al.
Published: (2025)
HumanSAM: Classifying Human-centric Forgery Videos in Human Spatial, Appearance, and Motion Anomaly
by: Liu, Chang, et al.
Published: (2025)
by: Liu, Chang, et al.
Published: (2025)
HMD-Poser: On-Device Real-time Human Motion Tracking from Scalable Sparse Observations
by: Dai, Peng, et al.
Published: (2024)
by: Dai, Peng, et al.
Published: (2024)
Spatial Degradation-Aware and Temporal Consistent Diffusion Model for Compressed Video Super-Resolution
by: An, Hongyu, et al.
Published: (2025)
by: An, Hongyu, et al.
Published: (2025)
TraceVision: Trajectory-Aware Vision-Language Model for Human-Like Spatial Understanding
by: Yang, Fan, et al.
Published: (2026)
by: Yang, Fan, et al.
Published: (2026)
Real-time Monocular Depth Estimation on Embedded Systems
by: Feng, Cheng, et al.
Published: (2023)
by: Feng, Cheng, et al.
Published: (2023)
A Spatial-Frequency Aware Multi-Scale Fusion Network for Real-Time Deepfake Detection
by: Lv, Libo, et al.
Published: (2025)
by: Lv, Libo, et al.
Published: (2025)
SPAD : Spatially Aware Multiview Diffusers
by: Kant, Yash, et al.
Published: (2024)
by: Kant, Yash, et al.
Published: (2024)
Ins-HOI: Instance Aware Human-Object Interactions Recovery
by: Zhang, Jiajun, et al.
Published: (2023)
by: Zhang, Jiajun, et al.
Published: (2023)
Boosting Quantitive and Spatial Awareness for Zero-Shot Object Counting
by: Zhang, Da, et al.
Published: (2026)
by: Zhang, Da, et al.
Published: (2026)
GenSpace: Benchmarking Spatially-Aware Image Generation
by: Wang, Zehan, et al.
Published: (2025)
by: Wang, Zehan, et al.
Published: (2025)
Refining CLIP's Spatial Awareness: A Visual-Centric Perspective
by: Qiu, Congpei, et al.
Published: (2025)
by: Qiu, Congpei, et al.
Published: (2025)
Similar Items
-
AV-Flow: Transforming Text to Audio-Visual Human-like Interactions
by: Chatziagapi, Aggelina, et al.
Published: (2025) -
From Audio to Photoreal Embodiment: Synthesizing Humans in Conversations
by: Ng, Evonne, et al.
Published: (2024) -
DuoMo: Dual Motion Diffusion for World-Space Human Reconstruction
by: Wang, Yufu, et al.
Published: (2026) -
Embody 3D: A Large-scale Multimodal Motion and Behavior Dataset
by: McLean, Claire, et al.
Published: (2025) -
Pose Priors from Language Models
by: Subramanian, Sanjay, et al.
Published: (2024)