Guardado en:
| Autores principales: | Ng, Evonne, Zhang, Siwei, Chen, Zhang, Zollhoefer, Michael, Richard, Alexander |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2602.18432 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
AV-Flow: Transforming Text to Audio-Visual Human-like Interactions
por: Chatziagapi, Aggelina, et al.
Publicado: (2025)
por: Chatziagapi, Aggelina, et al.
Publicado: (2025)
From Audio to Photoreal Embodiment: Synthesizing Humans in Conversations
por: Ng, Evonne, et al.
Publicado: (2024)
por: Ng, Evonne, et al.
Publicado: (2024)
DuoMo: Dual Motion Diffusion for World-Space Human Reconstruction
por: Wang, Yufu, et al.
Publicado: (2026)
por: Wang, Yufu, et al.
Publicado: (2026)
Embody 3D: A Large-scale Multimodal Motion and Behavior Dataset
por: McLean, Claire, et al.
Publicado: (2025)
por: McLean, Claire, et al.
Publicado: (2025)
Pose Priors from Language Models
por: Subramanian, Sanjay, et al.
Publicado: (2024)
por: Subramanian, Sanjay, et al.
Publicado: (2024)
Diffusion Forcing for Multi-Agent Interaction Sequence Modeling
por: Maluleke, Vongani H., et al.
Publicado: (2025)
por: Maluleke, Vongani H., et al.
Publicado: (2025)
RoHM: Robust Human Motion Reconstruction via Diffusion
por: Zhang, Siwei, et al.
Publicado: (2024)
por: Zhang, Siwei, et al.
Publicado: (2024)
3D Human Pose Estimation via Spatial Graph Order Attention and Temporal Body Aware Transformer
por: Aouaidjia, Kamel, et al.
Publicado: (2025)
por: Aouaidjia, Kamel, et al.
Publicado: (2025)
A Semantic-Aware and Multi-Guided Network for Infrared-Visible Image Fusion
por: Zhang, Xiaoli, et al.
Publicado: (2024)
por: Zhang, Xiaoli, et al.
Publicado: (2024)
Masked Modeling for Human Motion Recovery Under Occlusions
por: Qian, Zhiyin, et al.
Publicado: (2026)
por: Qian, Zhiyin, et al.
Publicado: (2026)
End-to-End Spatial-Temporal Transformer for Real-time 4D HOI Reconstruction
por: Zhang, Haoyu, et al.
Publicado: (2026)
por: Zhang, Haoyu, et al.
Publicado: (2026)
Spatial-Aware Self-Supervision for Medical 3D Imaging with Multi-Granularity Observable Tasks
por: Zhang, Yiqin, et al.
Publicado: (2025)
por: Zhang, Yiqin, et al.
Publicado: (2025)
PhyMAGIC: Physical Motion-Aware Generative Inference with Confidence-guided LLM
por: Meng, Siwei, et al.
Publicado: (2025)
por: Meng, Siwei, et al.
Publicado: (2025)
Human Gaussian Splatting: Real-time Rendering of Animatable Avatars
por: Moreau, Arthur, et al.
Publicado: (2023)
por: Moreau, Arthur, et al.
Publicado: (2023)
SPC-NeRF: Spatial Predictive Compression for Voxel Based Radiance Field
por: Song, Zetian, et al.
Publicado: (2024)
por: Song, Zetian, et al.
Publicado: (2024)
VTAgent: Agentic Keyframe Anchoring for Evidence-Aware Video TextVQA
por: He, Haibin, et al.
Publicado: (2026)
por: He, Haibin, et al.
Publicado: (2026)
Patho-AgenticRAG: Towards Multimodal Agentic Retrieval-Augmented Generation for Pathology VLMs via Reinforcement Learning
por: Zhang, Wenchuan, et al.
Publicado: (2025)
por: Zhang, Wenchuan, et al.
Publicado: (2025)
Real-time High-fidelity Gaussian Human Avatars with Position-based Interpolation of Spatially Distributed MLPs
por: Zhan, Youyi, et al.
Publicado: (2025)
por: Zhan, Youyi, et al.
Publicado: (2025)
Active Exploring like a Pigeon: Reinforcing Spatial Reasoning via Agentic Vision-Language Models
por: Deng, Wei, et al.
Publicado: (2026)
por: Deng, Wei, et al.
Publicado: (2026)
Real-time Spatial-temporal Traversability Assessment via Feature-based Sparse Gaussian Process
por: Hou, Zhenyu, et al.
Publicado: (2025)
por: Hou, Zhenyu, et al.
Publicado: (2025)
MoSA: Motion-Coherent Human Video Generation via Structure-Appearance Decoupling
por: Wang, Haoyu, et al.
Publicado: (2025)
por: Wang, Haoyu, et al.
Publicado: (2025)
Egocentric Visibility-Aware Human Pose Estimation
por: Dai, Peng, et al.
Publicado: (2026)
por: Dai, Peng, et al.
Publicado: (2026)
Fusing Pixels and Genes: Spatially-Aware Learning in Computational Pathology
por: Han, Minghao, et al.
Publicado: (2026)
por: Han, Minghao, et al.
Publicado: (2026)
TACR-YOLO: A Real-time Detection Framework for Abnormal Human Behaviors Enhanced with Coordinate and Task-Aware Representations
por: Yin, Xinyi, et al.
Publicado: (2025)
por: Yin, Xinyi, et al.
Publicado: (2025)
LGM-Pose: A Lightweight Global Modeling Network for Real-time Human Pose Estimation
por: Guo, Biao, et al.
Publicado: (2025)
por: Guo, Biao, et al.
Publicado: (2025)
EditRefiner: A Human-Aligned Agentic Framework for Image Editing Refinement
por: Xu, Zitong, et al.
Publicado: (2026)
por: Xu, Zitong, et al.
Publicado: (2026)
OARS: Process-Aware Online Alignment for Generative Real-World Image Super-Resolution
por: Zhao, Shijie, et al.
Publicado: (2026)
por: Zhao, Shijie, et al.
Publicado: (2026)
GaussianGAN: Real-Time Photorealistic controllable Human Avatars
por: Lakhal, Mohamed Ilyes, et al.
Publicado: (2025)
por: Lakhal, Mohamed Ilyes, et al.
Publicado: (2025)
PARSE: Part-Aware Relational Spatial Modeling
por: Bai, Yinuo, et al.
Publicado: (2026)
por: Bai, Yinuo, et al.
Publicado: (2026)
HumanSAM: Classifying Human-centric Forgery Videos in Human Spatial, Appearance, and Motion Anomaly
por: Liu, Chang, et al.
Publicado: (2025)
por: Liu, Chang, et al.
Publicado: (2025)
HMD-Poser: On-Device Real-time Human Motion Tracking from Scalable Sparse Observations
por: Dai, Peng, et al.
Publicado: (2024)
por: Dai, Peng, et al.
Publicado: (2024)
Real-time Monocular Depth Estimation on Embedded Systems
por: Feng, Cheng, et al.
Publicado: (2023)
por: Feng, Cheng, et al.
Publicado: (2023)
Spatial Degradation-Aware and Temporal Consistent Diffusion Model for Compressed Video Super-Resolution
por: An, Hongyu, et al.
Publicado: (2025)
por: An, Hongyu, et al.
Publicado: (2025)
SPAD : Spatially Aware Multiview Diffusers
por: Kant, Yash, et al.
Publicado: (2024)
por: Kant, Yash, et al.
Publicado: (2024)
A Spatial-Frequency Aware Multi-Scale Fusion Network for Real-Time Deepfake Detection
por: Lv, Libo, et al.
Publicado: (2025)
por: Lv, Libo, et al.
Publicado: (2025)
TraceVision: Trajectory-Aware Vision-Language Model for Human-Like Spatial Understanding
por: Yang, Fan, et al.
Publicado: (2026)
por: Yang, Fan, et al.
Publicado: (2026)
GenSpace: Benchmarking Spatially-Aware Image Generation
por: Wang, Zehan, et al.
Publicado: (2025)
por: Wang, Zehan, et al.
Publicado: (2025)
Boosting Quantitive and Spatial Awareness for Zero-Shot Object Counting
por: Zhang, Da, et al.
Publicado: (2026)
por: Zhang, Da, et al.
Publicado: (2026)
Refining CLIP's Spatial Awareness: A Visual-Centric Perspective
por: Qiu, Congpei, et al.
Publicado: (2025)
por: Qiu, Congpei, et al.
Publicado: (2025)
STAR: Spatial-Temporal Augmentation with Text-to-Video Models for Real-World Video Super-Resolution
por: Xie, Rui, et al.
Publicado: (2025)
por: Xie, Rui, et al.
Publicado: (2025)
Ejemplares similares
-
AV-Flow: Transforming Text to Audio-Visual Human-like Interactions
por: Chatziagapi, Aggelina, et al.
Publicado: (2025) -
From Audio to Photoreal Embodiment: Synthesizing Humans in Conversations
por: Ng, Evonne, et al.
Publicado: (2024) -
DuoMo: Dual Motion Diffusion for World-Space Human Reconstruction
por: Wang, Yufu, et al.
Publicado: (2026) -
Embody 3D: A Large-scale Multimodal Motion and Behavior Dataset
por: McLean, Claire, et al.
Publicado: (2025) -
Pose Priors from Language Models
por: Subramanian, Sanjay, et al.
Publicado: (2024)