Controllable Egocentric Video Generation via Occlusion-Aware Sparse 3D Hand Joints
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Chenyangguang, Ye, Botao, Chen, Boqi, Delitzas, Alexandros, Wang, Fangjinhua, Pollefeys, Marc, Wang, Xi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Open-Vocabulary Functional 3D Scene Graphs for Real-World Indoor Spaces
by: Zhang, Chenyangguang, et al.
Published: (2025)
by: Zhang, Chenyangguang, et al.
Published: (2025)
SIGHT: Synthesizing Image-Text Conditioned and Geometry-Guided 3D Hand-Object Trajectories
by: Gavryushin, Alexey, et al.
Published: (2025)
by: Gavryushin, Alexey, et al.
Published: (2025)
FunRec: Reconstructing Functional 3D Scenes from Egocentric Interaction Videos
by: Delitzas, Alexandros, et al.
Published: (2026)
by: Delitzas, Alexandros, et al.
Published: (2026)
Hierarchical and Holistic Open-Vocabulary Functional 3D Scene Graphs for Indoor Spaces
by: Hu, Xinggang, et al.
Published: (2026)
by: Hu, Xinggang, et al.
Published: (2026)
YoNoSplat: You Only Need One Model for Feedforward 3D Gaussian Splatting
by: Ye, Botao, et al.
Published: (2025)
by: Ye, Botao, et al.
Published: (2025)
REACT3D: Recovering Articulations for Interactive Physical 3D Scenes
by: Huang, Zhao, et al.
Published: (2025)
by: Huang, Zhao, et al.
Published: (2025)
Lightweight and Accurate Multi-View Stereo with Confidence-Aware Diffusion Model
by: Wang, Fangjinhua, et al.
Published: (2025)
by: Wang, Fangjinhua, et al.
Published: (2025)
MegaFlow: Zero-Shot Large Displacement Optical Flow
by: Zhang, Dingxi, et al.
Published: (2026)
by: Zhang, Dingxi, et al.
Published: (2026)
EgoGaussian: Dynamic Scene Understanding from Egocentric Video with 3D Gaussian Splatting
by: Zhang, Daiwei, et al.
Published: (2024)
by: Zhang, Daiwei, et al.
Published: (2024)
Synthesizing Consistent Novel Views via 3D Epipolar Attention without Re-Training
by: Ye, Botao, et al.
Published: (2025)
by: Ye, Botao, et al.
Published: (2025)
No Pose, No Problem in 4D: Feed-Forward Dynamic Gaussians from Unposed Multi-View Videos
by: Balice, Matteo, et al.
Published: (2026)
by: Balice, Matteo, et al.
Published: (2026)
UniSDF: Unifying Neural Representations for High-Fidelity 3D Reconstruction of Complex Scenes with Reflections
by: Wang, Fangjinhua, et al.
Published: (2023)
by: Wang, Fangjinhua, et al.
Published: (2023)
GLACE: Global Local Accelerated Coordinate Encoding
by: Wang, Fangjinhua, et al.
Published: (2024)
by: Wang, Fangjinhua, et al.
Published: (2024)
ImLoc: Revisiting Visual Localization with Image-based Representation
by: Jiang, Xudong, et al.
Published: (2026)
by: Jiang, Xudong, et al.
Published: (2026)
R-SCoRe: Revisiting Scene Coordinate Regression for Robust Large-Scale Visual Localization
by: Jiang, Xudong, et al.
Published: (2025)
by: Jiang, Xudong, et al.
Published: (2025)
No Pose, No Problem: Surprisingly Simple 3D Gaussian Splats from Sparse Unposed Images
by: Ye, Botao, et al.
Published: (2024)
by: Ye, Botao, et al.
Published: (2024)
MOHO: Learning Single-view Hand-held Object Reconstruction with Multi-view Occlusion-Aware Supervision
by: Zhang, Chenyangguang, et al.
Published: (2023)
by: Zhang, Chenyangguang, et al.
Published: (2023)
From Indoor to Open World: Revealing the Spatial Reasoning Gap in MLLMs
by: Wu, Mingrui, et al.
Published: (2025)
by: Wu, Mingrui, et al.
Published: (2025)
EgoPressure: A Dataset for Hand Pressure and Pose Estimation in Egocentric Vision
by: Zhao, Yiming, et al.
Published: (2024)
by: Zhao, Yiming, et al.
Published: (2024)
StableHand: Quality-Aware Flow Matching for World-Space Dual-Hand Motion Estimation from Egocentric Video
by: Zeng, Huajian, et al.
Published: (2026)
by: Zeng, Huajian, et al.
Published: (2026)
Interaction-Aware 4D Gaussian Splatting for Dynamic Hand-Object Interaction Reconstruction
by: Tian, Hao, et al.
Published: (2025)
by: Tian, Hao, et al.
Published: (2025)
HaWoR: World-Space Hand Motion Reconstruction from Egocentric Videos
by: Zhang, Jinglei, et al.
Published: (2025)
by: Zhang, Jinglei, et al.
Published: (2025)
MAPLE: Encoding Dexterous Robotic Manipulation Priors Learned From Egocentric Videos
by: Gavryushin, Alexey, et al.
Published: (2025)
by: Gavryushin, Alexey, et al.
Published: (2025)
VisualChef: Generating Visual Aids in Cooking via Mask Inpainting
by: Kuzyk, Oleh, et al.
Published: (2025)
by: Kuzyk, Oleh, et al.
Published: (2025)
MADiff: Motion-Aware Mamba Diffusion Models for Hand Trajectory Prediction on Egocentric Videos
by: Ma, Junyi, et al.
Published: (2024)
by: Ma, Junyi, et al.
Published: (2024)
Search3D: Hierarchical Open-Vocabulary 3D Segmentation
by: Takmaz, Ayca, et al.
Published: (2024)
by: Takmaz, Ayca, et al.
Published: (2024)
Lost & Found: Tracking Changes from Egocentric Observations in 3D Dynamic Scene Graphs
by: Behrens, Tjark, et al.
Published: (2024)
by: Behrens, Tjark, et al.
Published: (2024)
EMAG: Ego-motion Aware and Generalizable 2D Hand Forecasting from Egocentric Videos
by: Hatano, Masashi, et al.
Published: (2024)
by: Hatano, Masashi, et al.
Published: (2024)
EgoHandICL: Egocentric 3D Hand Reconstruction with In-Context Learning
by: Xie, Binzhu, et al.
Published: (2026)
by: Xie, Binzhu, et al.
Published: (2026)
HandGCAT: Occlusion-Robust 3D Hand Mesh Reconstruction from Monocular Images
by: Wang, Shuaibing, et al.
Published: (2024)
by: Wang, Shuaibing, et al.
Published: (2024)
TriSplat: Simulation-Ready Feed-Forward 3D Scene Reconstruction
by: Wang, Weijie, et al.
Published: (2026)
by: Wang, Weijie, et al.
Published: (2026)
Recognizing Hand Use and Hand Role at Home After Stroke from Egocentric Video
by: Tsai, Meng-Fen, et al.
Published: (2022)
by: Tsai, Meng-Fen, et al.
Published: (2022)
EgoControl: Controllable Egocentric Video Generation via 3D Full-Body Poses
by: Pallotta, Enrico, et al.
Published: (2025)
by: Pallotta, Enrico, et al.
Published: (2025)
DepthSplat: Connecting Gaussian Splatting and Depth
by: Xu, Haofei, et al.
Published: (2024)
by: Xu, Haofei, et al.
Published: (2024)
WHOLE: World-Grounded Hand-Object Lifted from Egocentric Videos
by: Ye, Yufei, et al.
Published: (2026)
by: Ye, Yufei, et al.
Published: (2026)
EgoGrasp: World-Space Hand-Object Interaction Estimation from Egocentric Videos
by: Fu, Hongming, et al.
Published: (2026)
by: Fu, Hongming, et al.
Published: (2026)
Hand2World: Autoregressive Egocentric Interaction Generation via Free-Space Hand Gestures
by: Wang, Yuxi, et al.
Published: (2026)
by: Wang, Yuxi, et al.
Published: (2026)
Occlusion-Aware 3D Hand-Object Pose Estimation with Masked AutoEncoders
by: Yang, Hui, et al.
Published: (2025)
by: Yang, Hui, et al.
Published: (2025)
Diff-IP2D: Diffusion-Based Hand-Object Interaction Prediction on Egocentric Videos
by: Ma, Junyi, et al.
Published: (2024)
by: Ma, Junyi, et al.
Published: (2024)
GenHOI: Generalized Hand-Object Pose Estimation with Occlusion Awareness
by: Yang, Hui, et al.
Published: (2026)
by: Yang, Hui, et al.
Published: (2026)
Similar Items
-
Open-Vocabulary Functional 3D Scene Graphs for Real-World Indoor Spaces
by: Zhang, Chenyangguang, et al.
Published: (2025) -
SIGHT: Synthesizing Image-Text Conditioned and Geometry-Guided 3D Hand-Object Trajectories
by: Gavryushin, Alexey, et al.
Published: (2025) -
FunRec: Reconstructing Functional 3D Scenes from Egocentric Interaction Videos
by: Delitzas, Alexandros, et al.
Published: (2026) -
Hierarchical and Holistic Open-Vocabulary Functional 3D Scene Graphs for Indoor Spaces
by: Hu, Xinggang, et al.
Published: (2026) -
YoNoSplat: You Only Need One Model for Feedforward 3D Gaussian Splatting
by: Ye, Botao, et al.
Published: (2025)