SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios
Fuente:
arXiv
Saved in:
| Main Authors: | Dang, Lingwei, Shao, Ruizhi, Zhang, Hongwen, Min, Wei, Liu, Yebin, Wu, Qingyao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SyncMV4D: Synchronized Multi-view Joint Diffusion of Appearance and Motion for Hand-Object Interaction Synthesis
by: Dang, Lingwei, et al.
Published: (2025)
by: Dang, Lingwei, et al.
Published: (2025)
ManiVideo: Generating Hand-Object Manipulation Video with Dexterous and Generalizable Grasping
by: Pang, Youxin, et al.
Published: (2024)
by: Pang, Youxin, et al.
Published: (2024)
HHMR: Holistic Hand Mesh Recovery by Enhancing the Multimodal Controllability of Graph Diffusion Models
by: Li, Mengcheng, et al.
Published: (2024)
by: Li, Mengcheng, et al.
Published: (2024)
UniMo: Unifying 2D Video and 3D Human Motion with an Autoregressive Framework
by: Pang, Youxin, et al.
Published: (2025)
by: Pang, Youxin, et al.
Published: (2025)
Human4DiT: 360-degree Human Video Generation with 4D Diffusion Transformer
by: Shao, Ruizhi, et al.
Published: (2024)
by: Shao, Ruizhi, et al.
Published: (2024)
DevilSight: Augmenting Monocular Human Avatar Reconstruction through a Virtual Perspective
by: Chen, Yushuo, et al.
Published: (2025)
by: Chen, Yushuo, et al.
Published: (2025)
Ins-HOI: Instance Aware Human-Object Interactions Recovery
by: Zhang, Jiajun, et al.
Published: (2023)
by: Zhang, Jiajun, et al.
Published: (2023)
HOSIG: Full-Body Human-Object-Scene Interaction Generation with Hierarchical Scene Perception
by: Yao, Wei, et al.
Published: (2025)
by: Yao, Wei, et al.
Published: (2025)
ManiDext: Hand-Object Manipulation Synthesis via Continuous Correspondence Embeddings and Residual-Guided Diffusion
by: Zhang, Jiajun, et al.
Published: (2024)
by: Zhang, Jiajun, et al.
Published: (2024)
SViTT-Ego: A Sparse Video-Text Transformer for Egocentric Video
by: Valdez, Hector A., et al.
Published: (2024)
by: Valdez, Hector A., et al.
Published: (2024)
Video Anomaly Detection with Semantics-Aware Information Bottleneck
by: Li, Juntong, et al.
Published: (2025)
by: Li, Juntong, et al.
Published: (2025)
DeCo: Decoupled Human-Centered Diffusion Video Editing with Motion Consistency
by: Zhong, Xiaojing, et al.
Published: (2024)
by: Zhong, Xiaojing, et al.
Published: (2024)
Not Blind but Silenced: Rebalancing Vision and Language via Adversarial Counter-Commonsense Equilibrium
by: Xiao, Qingxin, et al.
Published: (2026)
by: Xiao, Qingxin, et al.
Published: (2026)
Interspatial Attention for Efficient 4D Human Video Generation
by: Shao, Ruizhi, et al.
Published: (2025)
by: Shao, Ruizhi, et al.
Published: (2025)
Layered 3D Human Generation via Semantic-Aware Diffusion Model
by: Wang, Yi, et al.
Published: (2023)
by: Wang, Yi, et al.
Published: (2023)
CloSET: Modeling Clothed Humans on Continuous Surface with Explicit Template Decomposition
by: Zhang, Hongwen, et al.
Published: (2023)
by: Zhang, Hongwen, et al.
Published: (2023)
SViM3D: Stable Video Material Diffusion for Single Image 3D Generation
by: Engelhardt, Andreas, et al.
Published: (2025)
by: Engelhardt, Andreas, et al.
Published: (2025)
GPHM: Gaussian Parametric Head Model for Monocular Head Avatar Reconstruction
by: Xu, Yuelang, et al.
Published: (2024)
by: Xu, Yuelang, et al.
Published: (2024)
OmniHands: Towards Robust 4D Hand Mesh Recovery via A Versatile Transformer
by: Lin, Dixuan, et al.
Published: (2024)
by: Lin, Dixuan, et al.
Published: (2024)
Learning Explicit Contact for Implicit Reconstruction of Hand-held Objects from Monocular Images
by: Hu, Junxing, et al.
Published: (2023)
by: Hu, Junxing, et al.
Published: (2023)
GAF: Gaussian Action Field as a 4D Representation for Dynamic World Modeling in Robotic Manipulation
by: Chai, Ying, et al.
Published: (2025)
by: Chai, Ying, et al.
Published: (2025)
W-HMR: Monocular Human Mesh Recovery in World Space with Weak-Supervised Calibration
by: Yao, Wei, et al.
Published: (2023)
by: Yao, Wei, et al.
Published: (2023)
MEgoHand: Multimodal Egocentric Hand-Object Interaction Motion Generation
by: Zhou, Bohan, et al.
Published: (2025)
by: Zhou, Bohan, et al.
Published: (2025)
CoMo: Compositional Motion Customization for Text-to-Video Generation
by: Xu, Youcan, et al.
Published: (2025)
by: Xu, Youcan, et al.
Published: (2025)
MoVideo: Motion-Aware Video Generation with Diffusion Models
by: Liang, Jingyun, et al.
Published: (2023)
by: Liang, Jingyun, et al.
Published: (2023)
HandDiffuse: Generative Controllers for Two-Hand Interactions via Diffusion Models
by: Lin, Pei, et al.
Published: (2023)
by: Lin, Pei, et al.
Published: (2023)
Recovering 3D Human Mesh from Monocular Images: A Survey
by: Tian, Yating, et al.
Published: (2022)
by: Tian, Yating, et al.
Published: (2022)
HandX: Scaling Bimanual Motion and Interaction Generation
by: Zhang, Zimu, et al.
Published: (2026)
by: Zhang, Zimu, et al.
Published: (2026)
SegMo: Segment-aligned Text to 3D Human Motion Generation
by: Dang, Bowen, et al.
Published: (2025)
by: Dang, Bowen, et al.
Published: (2025)
PAD-Hand: Physics-Aware Diffusion for Hand Motion Recovery
by: Ismayilzada, Elkhan, et al.
Published: (2026)
by: Ismayilzada, Elkhan, et al.
Published: (2026)
DiMo: Discrete Diffusion Modeling for Motion Generation and Understanding
by: Zhang, Ning, et al.
Published: (2026)
by: Zhang, Ning, et al.
Published: (2026)
Syncphony: Synchronized Audio-to-Video Generation with Diffusion Transformers
by: Song, Jibin, et al.
Published: (2025)
by: Song, Jibin, et al.
Published: (2025)
MADiff: Motion-Aware Mamba Diffusion Models for Hand Trajectory Prediction on Egocentric Videos
by: Ma, Junyi, et al.
Published: (2024)
by: Ma, Junyi, et al.
Published: (2024)
MoAlign: Motion-Centric Representation Alignment for Video Diffusion Models
by: Bhowmik, Aritra, et al.
Published: (2025)
by: Bhowmik, Aritra, et al.
Published: (2025)
Mix3R: Mixing Feed-forward Reconstruction and Generative 3D Priors for Joint Multi-view Aligned 3D Reconstruction and Pose Estimation
by: Lin, Siyou, et al.
Published: (2026)
by: Lin, Siyou, et al.
Published: (2026)
InterHandGen: Two-Hand Interaction Generation via Cascaded Reverse Diffusion
by: Lee, Jihyun, et al.
Published: (2024)
by: Lee, Jihyun, et al.
Published: (2024)
SemanticSplat: Feed-Forward 3D Scene Understanding with Language-Aware Gaussian Fields
by: Li, Qijing, et al.
Published: (2025)
by: Li, Qijing, et al.
Published: (2025)
4DEquine: Disentangling Motion and Appearance for 4D Equine Reconstruction from Monocular Video
by: Lyu, Jin, et al.
Published: (2026)
by: Lyu, Jin, et al.
Published: (2026)
LaMoD: Latent Motion Diffusion Model For Myocardial Strain Generation
by: Xing, Jiarui, et al.
Published: (2024)
by: Xing, Jiarui, et al.
Published: (2024)
MoTrans: Customized Motion Transfer with Text-driven Video Diffusion Models
by: Li, Xiaomin, et al.
Published: (2024)
by: Li, Xiaomin, et al.
Published: (2024)
Similar Items
-
SyncMV4D: Synchronized Multi-view Joint Diffusion of Appearance and Motion for Hand-Object Interaction Synthesis
by: Dang, Lingwei, et al.
Published: (2025) -
ManiVideo: Generating Hand-Object Manipulation Video with Dexterous and Generalizable Grasping
by: Pang, Youxin, et al.
Published: (2024) -
HHMR: Holistic Hand Mesh Recovery by Enhancing the Multimodal Controllability of Graph Diffusion Models
by: Li, Mengcheng, et al.
Published: (2024) -
UniMo: Unifying 2D Video and 3D Human Motion with an Autoregressive Framework
by: Pang, Youxin, et al.
Published: (2025) -
Human4DiT: 360-degree Human Video Generation with 4D Diffusion Transformer
by: Shao, Ruizhi, et al.
Published: (2024)