DeVI: Physics-based Dexterous Human-Object Interaction via Synthetic Video Imitation
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Hyeonwoo, Kim, Jeonghwan, Cho, Kyungwon, Joo, Hanbyul |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ParaHome: Parameterizing Everyday Home Activities Towards 3D Generative Modeling of Human-Object Interactions
by: Kim, Jeonghwan, et al.
Published: (2024)
by: Kim, Jeonghwan, et al.
Published: (2024)
DAViD: Modeling Dynamic Affordance of 3D Objects Using Pre-trained Video Diffusion Models
by: Kim, Hyeonwoo, et al.
Published: (2025)
by: Kim, Hyeonwoo, et al.
Published: (2025)
OmniEgoCap: Camera-Agnostic Sequence-Level Egocentric Motion Reconstruction
by: Cho, Kyungwon, et al.
Published: (2025)
by: Cho, Kyungwon, et al.
Published: (2025)
Learning 3D Object Spatial Relationships from Pre-trained 2D Diffusion Models
by: Baik, Sangwon, et al.
Published: (2025)
by: Baik, Sangwon, et al.
Published: (2025)
Beyond the Contact: Discovering Comprehensive Affordance for 3D Objects from Pre-trained 2D Diffusion Models
by: Kim, Hyeonwoo, et al.
Published: (2024)
by: Kim, Hyeonwoo, et al.
Published: (2024)
Dexterous World Models
by: Kim, Byungjun, et al.
Published: (2025)
by: Kim, Byungjun, et al.
Published: (2025)
Target-Aware Video Diffusion Models
by: Kim, Taeksoo, et al.
Published: (2025)
by: Kim, Taeksoo, et al.
Published: (2025)
Vanast: Virtual Try-On with Human Image Animation via Synthetic Triplet Supervision
by: Cha, Hyunsoo, et al.
Published: (2026)
by: Cha, Hyunsoo, et al.
Published: (2026)
OmniRobotHome: A Multi-Camera Platform for Real-Time Multiadic Human-Robot Interaction
by: Lee, Junyoung, et al.
Published: (2026)
by: Lee, Junyoung, et al.
Published: (2026)
HRDexDB: A Large-Scale Dataset of Dexterous Human and Robotic Hand Grasps
by: Lim, Jongbin, et al.
Published: (2026)
by: Lim, Jongbin, et al.
Published: (2026)
Learning to Generate Human-Human-Object Interactions from Textual Descriptions
by: Na, Jeonghyeon, et al.
Published: (2025)
by: Na, Jeonghyeon, et al.
Published: (2025)
Text-Guided 6D Object Pose Rearrangement via Closed-Loop VLM Agents
by: Baik, Sangwon, et al.
Published: (2026)
by: Baik, Sangwon, et al.
Published: (2026)
UniSkill: Imitating Human Videos via Cross-Embodiment Skill Representations
by: Kim, Hanjung, et al.
Published: (2025)
by: Kim, Hanjung, et al.
Published: (2025)
GraspDiffusion: Synthesizing Realistic Whole-body Hand-Object Interaction
by: Kwon, Patrick, et al.
Published: (2024)
by: Kwon, Patrick, et al.
Published: (2024)
Durian: Dual Reference Image-Guided Portrait Animation with Attribute Transfer
by: Cha, Hyunsoo, et al.
Published: (2025)
by: Cha, Hyunsoo, et al.
Published: (2025)
Guess The Unseen: Dynamic 3D Scene Reconstruction from Partial 2D Glimpses
by: Lee, Inhee, et al.
Published: (2024)
by: Lee, Inhee, et al.
Published: (2024)
PEGASUS: Personalized Generative 3D Avatars with Composable Attributes
by: Cha, Hyunsoo, et al.
Published: (2024)
by: Cha, Hyunsoo, et al.
Published: (2024)
GALA: Generating Animatable Layered Assets from a Single Scan
by: Kim, Taeksoo, et al.
Published: (2024)
by: Kim, Taeksoo, et al.
Published: (2024)
Hierarchically Structured Neural Bones for Reconstructing Animatable Objects from Casual Videos
by: Jeon, Subin, et al.
Published: (2024)
by: Jeon, Subin, et al.
Published: (2024)
Locality-Aware Zero-Shot Human-Object Interaction Detection
by: Kim, Sanghyun, et al.
Published: (2025)
by: Kim, Sanghyun, et al.
Published: (2025)
Mocap Everyone Everywhere: Lightweight Motion Capture With Smartwatches and a Head-Mounted Camera
by: Lee, Jiye, et al.
Published: (2024)
by: Lee, Jiye, et al.
Published: (2024)
Improving Visual Token Reduction via Rectifying Distortions for Efficient Multimodal LLM Inference
by: Cho, Hyeonwoo, et al.
Published: (2026)
by: Cho, Hyeonwoo, et al.
Published: (2026)
PERSE: Personalized 3D Generative Avatars from A Single Portrait
by: Cha, Hyunsoo, et al.
Published: (2024)
by: Cha, Hyunsoo, et al.
Published: (2024)
Joint-Embedding Predictive Architecture for Self-Supervised Learning of Mask Classification Architecture
by: Kim, Dong-Hee, et al.
Published: (2024)
by: Kim, Dong-Hee, et al.
Published: (2024)
World Models for Learning Dexterous Hand-Object Interactions from Human Videos
by: Goswami, Raktim Gautam, et al.
Published: (2025)
by: Goswami, Raktim Gautam, et al.
Published: (2025)
CNG-SFDA:Clean-and-Noisy Region Guided Online-Offline Source-Free Domain Adaptation
by: Cho, Hyeonwoo, et al.
Published: (2024)
by: Cho, Hyeonwoo, et al.
Published: (2024)
MUST: Modality-Specific Representation-Aware Transformer for Diffusion-Enhanced Survival Prediction with Missing Modality
by: Kim, Kyungwon, et al.
Published: (2026)
by: Kim, Kyungwon, et al.
Published: (2026)
Multi-Granularity Video Object Segmentation
by: Lim, Sangbeom, et al.
Published: (2024)
by: Lim, Sangbeom, et al.
Published: (2024)
InterRVOS: Interaction-aware Referring Video Object Segmentation
by: Jin, Woojeong, et al.
Published: (2025)
by: Jin, Woojeong, et al.
Published: (2025)
Finer: Investigating and Enhancing Fine-Grained Visual Concept Recognition in Large Vision Language Models
by: Kim, Jeonghwan, et al.
Published: (2024)
by: Kim, Jeonghwan, et al.
Published: (2024)
Video Inference for Human Mesh Recovery with Vision Transformer
by: Cho, Hanbyel, et al.
Published: (2025)
by: Cho, Hanbyel, et al.
Published: (2025)
PLOT: Pseudo-Labeling via Video Object Tracking for Scalable Monocular 3D Object Detection
by: Lee, Seokyeong, et al.
Published: (2025)
by: Lee, Seokyeong, et al.
Published: (2025)
Search and Detect: Training-Free Long Tail Object Detection via Web-Image Retrieval
by: Sidhu, Mankeerat, et al.
Published: (2024)
by: Sidhu, Mankeerat, et al.
Published: (2024)
CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation
by: Luo, Xiangyang, et al.
Published: (2026)
by: Luo, Xiangyang, et al.
Published: (2026)
Decoupled Generative Modeling for Human-Object Interaction Synthesis
by: Jung, Hwanhee, et al.
Published: (2025)
by: Jung, Hwanhee, et al.
Published: (2025)
Robust Multimodal 3D Object Detection via Modality-Agnostic Decoding and Proximity-based Modality Ensemble
by: Cha, Juhan, et al.
Published: (2024)
by: Cha, Juhan, et al.
Published: (2024)
Dexterous Manipulation Policies from RGB Human Videos via 3D Hand-Object Trajectory Reconstruction
by: Chen, Hongyi, et al.
Published: (2026)
by: Chen, Hongyi, et al.
Published: (2026)
HairCUP: Hair Compositional Universal Prior for 3D Gaussian Avatars
by: Kim, Byungjun, et al.
Published: (2025)
by: Kim, Byungjun, et al.
Published: (2025)
ManiVideo: Generating Hand-Object Manipulation Video with Dexterous and Generalizable Grasping
by: Pang, Youxin, et al.
Published: (2024)
by: Pang, Youxin, et al.
Published: (2024)
WebSpline: Structure-Informed Splines for Real-Time 3D Gaussians from Monocular Videos
by: Park, Jongmin, et al.
Published: (2026)
by: Park, Jongmin, et al.
Published: (2026)
Similar Items
-
ParaHome: Parameterizing Everyday Home Activities Towards 3D Generative Modeling of Human-Object Interactions
by: Kim, Jeonghwan, et al.
Published: (2024) -
DAViD: Modeling Dynamic Affordance of 3D Objects Using Pre-trained Video Diffusion Models
by: Kim, Hyeonwoo, et al.
Published: (2025) -
OmniEgoCap: Camera-Agnostic Sequence-Level Egocentric Motion Reconstruction
by: Cho, Kyungwon, et al.
Published: (2025) -
Learning 3D Object Spatial Relationships from Pre-trained 2D Diffusion Models
by: Baik, Sangwon, et al.
Published: (2025) -
Beyond the Contact: Discovering Comprehensive Affordance for 3D Objects from Pre-trained 2D Diffusion Models
by: Kim, Hyeonwoo, et al.
Published: (2024)