Gespeichert in:
| Hauptverfasser: | Ng, Evonne, Zhang, Siwei, Chen, Zhang, Zollhoefer, Michael, Richard, Alexander |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2602.18432 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AV-Flow: Transforming Text to Audio-Visual Human-like Interactions
von: Chatziagapi, Aggelina, et al.
Veröffentlicht: (2025)
von: Chatziagapi, Aggelina, et al.
Veröffentlicht: (2025)
From Audio to Photoreal Embodiment: Synthesizing Humans in Conversations
von: Ng, Evonne, et al.
Veröffentlicht: (2024)
von: Ng, Evonne, et al.
Veröffentlicht: (2024)
DuoMo: Dual Motion Diffusion for World-Space Human Reconstruction
von: Wang, Yufu, et al.
Veröffentlicht: (2026)
von: Wang, Yufu, et al.
Veröffentlicht: (2026)
Embody 3D: A Large-scale Multimodal Motion and Behavior Dataset
von: McLean, Claire, et al.
Veröffentlicht: (2025)
von: McLean, Claire, et al.
Veröffentlicht: (2025)
Pose Priors from Language Models
von: Subramanian, Sanjay, et al.
Veröffentlicht: (2024)
von: Subramanian, Sanjay, et al.
Veröffentlicht: (2024)
Diffusion Forcing for Multi-Agent Interaction Sequence Modeling
von: Maluleke, Vongani H., et al.
Veröffentlicht: (2025)
von: Maluleke, Vongani H., et al.
Veröffentlicht: (2025)
RoHM: Robust Human Motion Reconstruction via Diffusion
von: Zhang, Siwei, et al.
Veröffentlicht: (2024)
von: Zhang, Siwei, et al.
Veröffentlicht: (2024)
3D Human Pose Estimation via Spatial Graph Order Attention and Temporal Body Aware Transformer
von: Aouaidjia, Kamel, et al.
Veröffentlicht: (2025)
von: Aouaidjia, Kamel, et al.
Veröffentlicht: (2025)
A Semantic-Aware and Multi-Guided Network for Infrared-Visible Image Fusion
von: Zhang, Xiaoli, et al.
Veröffentlicht: (2024)
von: Zhang, Xiaoli, et al.
Veröffentlicht: (2024)
Masked Modeling for Human Motion Recovery Under Occlusions
von: Qian, Zhiyin, et al.
Veröffentlicht: (2026)
von: Qian, Zhiyin, et al.
Veröffentlicht: (2026)
End-to-End Spatial-Temporal Transformer for Real-time 4D HOI Reconstruction
von: Zhang, Haoyu, et al.
Veröffentlicht: (2026)
von: Zhang, Haoyu, et al.
Veröffentlicht: (2026)
Spatial-Aware Self-Supervision for Medical 3D Imaging with Multi-Granularity Observable Tasks
von: Zhang, Yiqin, et al.
Veröffentlicht: (2025)
von: Zhang, Yiqin, et al.
Veröffentlicht: (2025)
PhyMAGIC: Physical Motion-Aware Generative Inference with Confidence-guided LLM
von: Meng, Siwei, et al.
Veröffentlicht: (2025)
von: Meng, Siwei, et al.
Veröffentlicht: (2025)
Human Gaussian Splatting: Real-time Rendering of Animatable Avatars
von: Moreau, Arthur, et al.
Veröffentlicht: (2023)
von: Moreau, Arthur, et al.
Veröffentlicht: (2023)
SPC-NeRF: Spatial Predictive Compression for Voxel Based Radiance Field
von: Song, Zetian, et al.
Veröffentlicht: (2024)
von: Song, Zetian, et al.
Veröffentlicht: (2024)
VTAgent: Agentic Keyframe Anchoring for Evidence-Aware Video TextVQA
von: He, Haibin, et al.
Veröffentlicht: (2026)
von: He, Haibin, et al.
Veröffentlicht: (2026)
Patho-AgenticRAG: Towards Multimodal Agentic Retrieval-Augmented Generation for Pathology VLMs via Reinforcement Learning
von: Zhang, Wenchuan, et al.
Veröffentlicht: (2025)
von: Zhang, Wenchuan, et al.
Veröffentlicht: (2025)
Real-time High-fidelity Gaussian Human Avatars with Position-based Interpolation of Spatially Distributed MLPs
von: Zhan, Youyi, et al.
Veröffentlicht: (2025)
von: Zhan, Youyi, et al.
Veröffentlicht: (2025)
Active Exploring like a Pigeon: Reinforcing Spatial Reasoning via Agentic Vision-Language Models
von: Deng, Wei, et al.
Veröffentlicht: (2026)
von: Deng, Wei, et al.
Veröffentlicht: (2026)
Real-time Spatial-temporal Traversability Assessment via Feature-based Sparse Gaussian Process
von: Hou, Zhenyu, et al.
Veröffentlicht: (2025)
von: Hou, Zhenyu, et al.
Veröffentlicht: (2025)
MoSA: Motion-Coherent Human Video Generation via Structure-Appearance Decoupling
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
Egocentric Visibility-Aware Human Pose Estimation
von: Dai, Peng, et al.
Veröffentlicht: (2026)
von: Dai, Peng, et al.
Veröffentlicht: (2026)
Fusing Pixels and Genes: Spatially-Aware Learning in Computational Pathology
von: Han, Minghao, et al.
Veröffentlicht: (2026)
von: Han, Minghao, et al.
Veröffentlicht: (2026)
TACR-YOLO: A Real-time Detection Framework for Abnormal Human Behaviors Enhanced with Coordinate and Task-Aware Representations
von: Yin, Xinyi, et al.
Veröffentlicht: (2025)
von: Yin, Xinyi, et al.
Veröffentlicht: (2025)
LGM-Pose: A Lightweight Global Modeling Network for Real-time Human Pose Estimation
von: Guo, Biao, et al.
Veröffentlicht: (2025)
von: Guo, Biao, et al.
Veröffentlicht: (2025)
EditRefiner: A Human-Aligned Agentic Framework for Image Editing Refinement
von: Xu, Zitong, et al.
Veröffentlicht: (2026)
von: Xu, Zitong, et al.
Veröffentlicht: (2026)
OARS: Process-Aware Online Alignment for Generative Real-World Image Super-Resolution
von: Zhao, Shijie, et al.
Veröffentlicht: (2026)
von: Zhao, Shijie, et al.
Veröffentlicht: (2026)
GaussianGAN: Real-Time Photorealistic controllable Human Avatars
von: Lakhal, Mohamed Ilyes, et al.
Veröffentlicht: (2025)
von: Lakhal, Mohamed Ilyes, et al.
Veröffentlicht: (2025)
PARSE: Part-Aware Relational Spatial Modeling
von: Bai, Yinuo, et al.
Veröffentlicht: (2026)
von: Bai, Yinuo, et al.
Veröffentlicht: (2026)
HumanSAM: Classifying Human-centric Forgery Videos in Human Spatial, Appearance, and Motion Anomaly
von: Liu, Chang, et al.
Veröffentlicht: (2025)
von: Liu, Chang, et al.
Veröffentlicht: (2025)
HMD-Poser: On-Device Real-time Human Motion Tracking from Scalable Sparse Observations
von: Dai, Peng, et al.
Veröffentlicht: (2024)
von: Dai, Peng, et al.
Veröffentlicht: (2024)
Real-time Monocular Depth Estimation on Embedded Systems
von: Feng, Cheng, et al.
Veröffentlicht: (2023)
von: Feng, Cheng, et al.
Veröffentlicht: (2023)
Spatial Degradation-Aware and Temporal Consistent Diffusion Model for Compressed Video Super-Resolution
von: An, Hongyu, et al.
Veröffentlicht: (2025)
von: An, Hongyu, et al.
Veröffentlicht: (2025)
SPAD : Spatially Aware Multiview Diffusers
von: Kant, Yash, et al.
Veröffentlicht: (2024)
von: Kant, Yash, et al.
Veröffentlicht: (2024)
A Spatial-Frequency Aware Multi-Scale Fusion Network for Real-Time Deepfake Detection
von: Lv, Libo, et al.
Veröffentlicht: (2025)
von: Lv, Libo, et al.
Veröffentlicht: (2025)
TraceVision: Trajectory-Aware Vision-Language Model for Human-Like Spatial Understanding
von: Yang, Fan, et al.
Veröffentlicht: (2026)
von: Yang, Fan, et al.
Veröffentlicht: (2026)
GenSpace: Benchmarking Spatially-Aware Image Generation
von: Wang, Zehan, et al.
Veröffentlicht: (2025)
von: Wang, Zehan, et al.
Veröffentlicht: (2025)
Boosting Quantitive and Spatial Awareness for Zero-Shot Object Counting
von: Zhang, Da, et al.
Veröffentlicht: (2026)
von: Zhang, Da, et al.
Veröffentlicht: (2026)
Refining CLIP's Spatial Awareness: A Visual-Centric Perspective
von: Qiu, Congpei, et al.
Veröffentlicht: (2025)
von: Qiu, Congpei, et al.
Veröffentlicht: (2025)
STAR: Spatial-Temporal Augmentation with Text-to-Video Models for Real-World Video Super-Resolution
von: Xie, Rui, et al.
Veröffentlicht: (2025)
von: Xie, Rui, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
AV-Flow: Transforming Text to Audio-Visual Human-like Interactions
von: Chatziagapi, Aggelina, et al.
Veröffentlicht: (2025) -
From Audio to Photoreal Embodiment: Synthesizing Humans in Conversations
von: Ng, Evonne, et al.
Veröffentlicht: (2024) -
DuoMo: Dual Motion Diffusion for World-Space Human Reconstruction
von: Wang, Yufu, et al.
Veröffentlicht: (2026) -
Embody 3D: A Large-scale Multimodal Motion and Behavior Dataset
von: McLean, Claire, et al.
Veröffentlicht: (2025) -
Pose Priors from Language Models
von: Subramanian, Sanjay, et al.
Veröffentlicht: (2024)