End-to-End Spatial-Temporal Transformer for Real-time 4D HOI Reconstruction
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Haoyu, Zhai, Wei, Yang, Yuhang, Cao, Yang, Zha, Zheng-Jun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
End-to-End HOI Reconstruction Transformer with Graph-based Encoding
von: Wang, Zhenrong, et al.
Veröffentlicht: (2025)
von: Wang, Zhenrong, et al.
Veröffentlicht: (2025)
LEMON: Learning 3D Human-Object Interaction Relation from 2D Images
von: Yang, Yuhang, et al.
Veröffentlicht: (2023)
von: Yang, Yuhang, et al.
Veröffentlicht: (2023)
HERO: Human Reaction Generation from Videos
von: Yu, Chengjun, et al.
Veröffentlicht: (2025)
von: Yu, Chengjun, et al.
Veröffentlicht: (2025)
TOUCH: Text-guided Controllable Generation of Free-Form Hand-Object Interactions
von: Han, Guangyi, et al.
Veröffentlicht: (2025)
von: Han, Guangyi, et al.
Veröffentlicht: (2025)
EgoChoir: Capturing 3D Human-Object Interaction Regions from Egocentric Views
von: Yang, Yuhang, et al.
Veröffentlicht: (2024)
von: Yang, Yuhang, et al.
Veröffentlicht: (2024)
GRACE: Estimating Geometry-level 3D Human-Scene Contact from 2D Images
von: Wang, Chengfeng, et al.
Veröffentlicht: (2025)
von: Wang, Chengfeng, et al.
Veröffentlicht: (2025)
GREAT: Geometry-Intention Collaborative Inference for Open-Vocabulary 3D Object Affordance Grounding
von: Shao, Yawen, et al.
Veröffentlicht: (2024)
von: Shao, Yawen, et al.
Veröffentlicht: (2024)
RAIN: Real-time Animation of Infinite Video Stream
von: Shu, Zhilei, et al.
Veröffentlicht: (2024)
von: Shu, Zhilei, et al.
Veröffentlicht: (2024)
Grounding 3D Scene Affordance From Egocentric Interactions
von: Liu, Cuiyu, et al.
Veröffentlicht: (2024)
von: Liu, Cuiyu, et al.
Veröffentlicht: (2024)
MATE: Motion-Augmented Temporal Consistency for Event-based Point Tracking
von: Han, Han, et al.
Veröffentlicht: (2024)
von: Han, Han, et al.
Veröffentlicht: (2024)
World4Drive: End-to-End Autonomous Driving via Intention-aware Physical Latent World Model
von: Zheng, Yupeng, et al.
Veröffentlicht: (2025)
von: Zheng, Yupeng, et al.
Veröffentlicht: (2025)
Gloria: Consistent Character Video Generation via Content Anchors
von: Yang, Yuhang, et al.
Veröffentlicht: (2026)
von: Yang, Yuhang, et al.
Veröffentlicht: (2026)
Real2Sim in HOI: Toward Physically Plausible HOI Reconstruction from Monocular Videos
von: Zhao, Yubo, et al.
Veröffentlicht: (2026)
von: Zhao, Yubo, et al.
Veröffentlicht: (2026)
End to End Face Reconstruction via Differentiable PnP
von: Lu, Yiren, et al.
Veröffentlicht: (2024)
von: Lu, Yiren, et al.
Veröffentlicht: (2024)
HFGS: 4D Gaussian Splatting with Emphasis on Spatial and Temporal High-Frequency Components for Endoscopic Scene Reconstruction
von: Zhao, Haoyu, et al.
Veröffentlicht: (2024)
von: Zhao, Haoyu, et al.
Veröffentlicht: (2024)
VisionSelector: End-to-End Learnable Visual Token Compression for Efficient Multimodal LLMs
von: Zhu, Jiaying, et al.
Veröffentlicht: (2025)
von: Zhu, Jiaying, et al.
Veröffentlicht: (2025)
EMoTive: Event-guided Trajectory Modeling for 3D Motion Estimation
von: Wan, Zengyu, et al.
Veröffentlicht: (2025)
von: Wan, Zengyu, et al.
Veröffentlicht: (2025)
ArtHOI: Articulated Human-Object Interaction Synthesis by 4D Reconstruction from Video Priors
von: Huang, Zihao, et al.
Veröffentlicht: (2026)
von: Huang, Zihao, et al.
Veröffentlicht: (2026)
Rethinking the Spatio-Temporal Alignment of End-to-End 3D Perception
von: Li, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Li, Xiaoyu, et al.
Veröffentlicht: (2025)
Event Stream Filtering via Probability Flux Estimation
von: Chen, Jinze, et al.
Veröffentlicht: (2025)
von: Chen, Jinze, et al.
Veröffentlicht: (2025)
Visual-Geometric Collaborative Guidance for Affordance Learning
von: Luo, Hongchen, et al.
Veröffentlicht: (2024)
von: Luo, Hongchen, et al.
Veröffentlicht: (2024)
Leverage Task Context for Object Affordance Ranking
von: Huang, Haojie, et al.
Veröffentlicht: (2024)
von: Huang, Haojie, et al.
Veröffentlicht: (2024)
DREAM: Document Reconstruction via End-to-end Autoregressive Model
von: Li, Xin, et al.
Veröffentlicht: (2025)
von: Li, Xin, et al.
Veröffentlicht: (2025)
EF-3DGS: Event-Aided Free-Trajectory 3D Gaussian Splatting
von: Liao, Bohao, et al.
Veröffentlicht: (2024)
von: Liao, Bohao, et al.
Veröffentlicht: (2024)
End-to-End Multi-Person Pose Estimation with Pose-Aware Video Transformer
von: Yu, Yonghui, et al.
Veröffentlicht: (2025)
von: Yu, Yonghui, et al.
Veröffentlicht: (2025)
Event-based Asynchronous HDR Imaging by Temporal Incident Light Modulation
von: Wu, Yuliang, et al.
Veröffentlicht: (2024)
von: Wu, Yuliang, et al.
Veröffentlicht: (2024)
End-to-End 3D Spatiotemporal Perception with Multimodal Fusion and V2X Collaboration
von: Yang, Zhenwei, et al.
Veröffentlicht: (2025)
von: Yang, Zhenwei, et al.
Veröffentlicht: (2025)
SIGMAN:Scaling 3D Human Gaussian Generation with Millions of Assets
von: Yang, Yuhang, et al.
Veröffentlicht: (2025)
von: Yang, Yuhang, et al.
Veröffentlicht: (2025)
Puzzles: Unbounded Video-Depth Augmentation for Scalable End-to-End 3D Reconstruction
von: Ma, Jiahao, et al.
Veröffentlicht: (2025)
von: Ma, Jiahao, et al.
Veröffentlicht: (2025)
End-to-End Training for Autoregressive Video Diffusion via Self-Resampling
von: Guo, Yuwei, et al.
Veröffentlicht: (2025)
von: Guo, Yuwei, et al.
Veröffentlicht: (2025)
End-to-End Streaming Video Temporal Action Segmentation with Reinforce Learning
von: Zhang, Jinrong, et al.
Veröffentlicht: (2023)
von: Zhang, Jinrong, et al.
Veröffentlicht: (2023)
End-to-End 4D Heart Mesh Recovery Across Full-Stack and Sparse Cardiac MRI
von: Chen, Yihong, et al.
Veröffentlicht: (2025)
von: Chen, Yihong, et al.
Veröffentlicht: (2025)
SimFlow: Simplified and End-to-End Training of Latent Normalizing Flows
von: Zhao, Qinyu, et al.
Veröffentlicht: (2025)
von: Zhao, Qinyu, et al.
Veröffentlicht: (2025)
FastDriveVLA: Efficient End-to-End Driving via Plug-and-Play Reconstruction-based Token Pruning
von: Cao, Jiajun, et al.
Veröffentlicht: (2025)
von: Cao, Jiajun, et al.
Veröffentlicht: (2025)
ResWorld: Temporal Residual World Model for End-to-End Autonomous Driving
von: Zhang, Jinqing, et al.
Veröffentlicht: (2026)
von: Zhang, Jinqing, et al.
Veröffentlicht: (2026)
End4: End-to-end Denoising Diffusion for Diffusion-Based Inpainting Detection
von: Wang, Fei, et al.
Veröffentlicht: (2025)
von: Wang, Fei, et al.
Veröffentlicht: (2025)
Semantic Segmentation and Scene Reconstruction of RGB-D Image Frames: An End-to-End Modular Pipeline for Robotic Applications
von: Zheng, Zhiwu, et al.
Veröffentlicht: (2024)
von: Zheng, Zhiwu, et al.
Veröffentlicht: (2024)
Event-based Visual Deformation Measurement
von: Wu, Yuliang, et al.
Veröffentlicht: (2026)
von: Wu, Yuliang, et al.
Veröffentlicht: (2026)
End-to-End Navigation with Vision Language Models: Transforming Spatial Reasoning into Question-Answering
von: Goetting, Dylan, et al.
Veröffentlicht: (2024)
von: Goetting, Dylan, et al.
Veröffentlicht: (2024)
DLAFormer: An End-to-End Transformer For Document Layout Analysis
von: Wang, Jiawei, et al.
Veröffentlicht: (2024)
von: Wang, Jiawei, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
End-to-End HOI Reconstruction Transformer with Graph-based Encoding
von: Wang, Zhenrong, et al.
Veröffentlicht: (2025) -
LEMON: Learning 3D Human-Object Interaction Relation from 2D Images
von: Yang, Yuhang, et al.
Veröffentlicht: (2023) -
HERO: Human Reaction Generation from Videos
von: Yu, Chengjun, et al.
Veröffentlicht: (2025) -
TOUCH: Text-guided Controllable Generation of Free-Form Hand-Object Interactions
von: Han, Guangyi, et al.
Veröffentlicht: (2025) -
EgoChoir: Capturing 3D Human-Object Interaction Regions from Egocentric Views
von: Yang, Yuhang, et al.
Veröffentlicht: (2024)