Controlling the World by Sleight of Hand
Fuente:
arXiv
Saved in:
| Main Authors: | Sudhakar, Sruthi, Liu, Ruoshi, Van Hoorick, Basile, Vondrick, Carl, Zemel, Richard |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dreamitate: Real-World Visuomotor Policy Learning via Video Generation
by: Liang, Junbang, et al.
Published: (2024)
by: Liang, Junbang, et al.
Published: (2024)
Whiteboard-of-Thought: Thinking Step-by-Step Across Modalities
by: Menon, Sachit, et al.
Published: (2024)
by: Menon, Sachit, et al.
Published: (2024)
Generative Camera Dolly: Extreme Monocular Dynamic Novel View Synthesis
by: Van Hoorick, Basile, et al.
Published: (2024)
by: Van Hoorick, Basile, et al.
Published: (2024)
Differentiable Robot Rendering
by: Liu, Ruoshi, et al.
Published: (2024)
by: Liu, Ruoshi, et al.
Published: (2024)
EraseDraw: Learning to Draw Step-by-Step via Erasing Objects from Images
by: Canberk, Alper, et al.
Published: (2024)
by: Canberk, Alper, et al.
Published: (2024)
Sin3DM: Learning a Diffusion Model from a Single 3D Textured Shape
by: Wu, Rundi, et al.
Published: (2023)
by: Wu, Rundi, et al.
Published: (2023)
SIRE: SE(3) Intrinsic Rigidity Embeddings
by: Smith, Cameron, et al.
Published: (2025)
by: Smith, Cameron, et al.
Published: (2025)
Fiducial Exoskeletons: Image-Centric Robot State Estimation
by: Smith, Cameron, et al.
Published: (2026)
by: Smith, Cameron, et al.
Published: (2026)
pix2gestalt: Amodal Segmentation by Synthesizing Wholes
by: Ozguroglu, Ege, et al.
Published: (2024)
by: Ozguroglu, Ege, et al.
Published: (2024)
4D Gaussian Splatting as a Learned Dynamical System
by: Asiimwe, Arnold Caleb, et al.
Published: (2025)
by: Asiimwe, Arnold Caleb, et al.
Published: (2025)
RoboDream: Compositional World Models for Scalable Robot Data Synthesis
by: Ye, Junjie, et al.
Published: (2026)
by: Ye, Junjie, et al.
Published: (2026)
AnchorDream: Repurposing Video Diffusion for Embodiment-Aware Robot Data Synthesis
by: Ye, Junjie, et al.
Published: (2025)
by: Ye, Junjie, et al.
Published: (2025)
GES: Generalized Exponential Splatting for Efficient Radiance Field Rendering
by: Hamdi, Abdullah, et al.
Published: (2024)
by: Hamdi, Abdullah, et al.
Published: (2024)
How Video Meetings Change Your Expression
by: Sarin, Sumit, et al.
Published: (2024)
by: Sarin, Sumit, et al.
Published: (2024)
Evolving Interpretable Visual Classifiers with Large Language Models
by: Chiquier, Mia, et al.
Published: (2024)
by: Chiquier, Mia, et al.
Published: (2024)
MedAutoCorrect: Image-Conditioned Autocorrection in Medical Reporting
by: Asiimwe, Arnold Caleb, et al.
Published: (2024)
by: Asiimwe, Arnold Caleb, et al.
Published: (2024)
New York Smells: A Large Multimodal Dataset for Olfaction
by: Ozguroglu, Ege, et al.
Published: (2025)
by: Ozguroglu, Ege, et al.
Published: (2025)
Teaching Humans Subtle Differences with DIFFusion
by: Chiquier, Mia, et al.
Published: (2025)
by: Chiquier, Mia, et al.
Published: (2025)
Two Hands Are Better Than One: Resolving Hand to Hand Intersections via Occupancy Networks
by: Ivashechkin, Maksym, et al.
Published: (2024)
by: Ivashechkin, Maksym, et al.
Published: (2024)
HandDiffuse: Generative Controllers for Two-Hand Interactions via Diffusion Models
by: Lin, Pei, et al.
Published: (2023)
by: Lin, Pei, et al.
Published: (2023)
Integrating Present and Past in Unsupervised Continual Learning
by: Zhang, Yipeng, et al.
Published: (2024)
by: Zhang, Yipeng, et al.
Published: (2024)
WorldGenBench: A World-Knowledge-Integrated Benchmark for Reasoning-Driven Text-to-Image Generation
by: Zhang, Daoan, et al.
Published: (2025)
by: Zhang, Daoan, et al.
Published: (2025)
HandOcc: NeRF-based Hand Rendering with Occupancy Networks
by: Ivashechkin, Maksym, et al.
Published: (2025)
by: Ivashechkin, Maksym, et al.
Published: (2025)
WHOLE: World-Grounded Hand-Object Lifted from Egocentric Videos
by: Ye, Yufei, et al.
Published: (2026)
by: Ye, Yufei, et al.
Published: (2026)
CAViAR: Critic-Augmented Video Agentic Reasoning
by: Menon, Sachit, et al.
Published: (2025)
by: Menon, Sachit, et al.
Published: (2025)
AnyView: Synthesizing Any Novel View in Dynamic Scenes
by: Van Hoorick, Basile, et al.
Published: (2026)
by: Van Hoorick, Basile, et al.
Published: (2026)
Hand2World: Autoregressive Egocentric Interaction Generation via Free-Space Hand Gestures
by: Wang, Yuxi, et al.
Published: (2026)
by: Wang, Yuxi, et al.
Published: (2026)
HandDiff: 3D Hand Pose Estimation with Diffusion on Image-Point Cloud
by: Cheng, Wencan, et al.
Published: (2024)
by: Cheng, Wencan, et al.
Published: (2024)
Generated Reality: Human-centric World Simulation using Interactive Video Generation with Hand and Camera Control
by: Xie, Linxi, et al.
Published: (2026)
by: Xie, Linxi, et al.
Published: (2026)
Egocentric World Model for Photorealistic Hand-Object Interaction Synthesis
by: Li, Dayou, et al.
Published: (2026)
by: Li, Dayou, et al.
Published: (2026)
DiSciPLE: Learning Interpretable Programs for Scientific Visual Discovery
by: Mall, Utkarsh, et al.
Published: (2025)
by: Mall, Utkarsh, et al.
Published: (2025)
FoundHand: Large-Scale Domain-Specific Learning for Controllable Hand Image Generation
by: Chen, Kefan, et al.
Published: (2024)
by: Chen, Kefan, et al.
Published: (2024)
UniHand: A Unified Model for Diverse Controlled 4D Hand Motion Modeling
by: Sun, Zhihao, et al.
Published: (2026)
by: Sun, Zhihao, et al.
Published: (2026)
Reinforcing 3D Understanding in Point-VLMs via Geometric Reward Credit Assignment
by: Chen, Jingkun, et al.
Published: (2026)
by: Chen, Jingkun, et al.
Published: (2026)
EgoSpot:Egocentric Multimodal Control for Hands-Free Mobile Manipulation
by: Zhang, Ganlin, et al.
Published: (2023)
by: Zhang, Ganlin, et al.
Published: (2023)
Do multimodal models imagine electric sheep?
by: Ramakrishnan, Santhosh Kumar, et al.
Published: (2026)
by: Ramakrishnan, Santhosh Kumar, et al.
Published: (2026)
SesaHand: Enhancing 3D Hand Reconstruction via Controllable Generation with Semantic and Structural Alignment
by: Zhao, Zhuoran, et al.
Published: (2026)
by: Zhao, Zhuoran, et al.
Published: (2026)
HandBooster: Boosting 3D Hand-Mesh Reconstruction by Conditional Synthesis and Sampling of Hand-Object Interactions
by: Xu, Hao, et al.
Published: (2024)
by: Xu, Hao, et al.
Published: (2024)
CompressNAS : A Fast and Efficient Technique for Model Compression using Decomposition
by: Sah, Sudhakar, et al.
Published: (2025)
by: Sah, Sudhakar, et al.
Published: (2025)
HandCraft: Anatomically Correct Restoration of Malformed Hands in Diffusion Generated Images
by: Qin, Zhenyue, et al.
Published: (2024)
by: Qin, Zhenyue, et al.
Published: (2024)
Similar Items
-
Dreamitate: Real-World Visuomotor Policy Learning via Video Generation
by: Liang, Junbang, et al.
Published: (2024) -
Whiteboard-of-Thought: Thinking Step-by-Step Across Modalities
by: Menon, Sachit, et al.
Published: (2024) -
Generative Camera Dolly: Extreme Monocular Dynamic Novel View Synthesis
by: Van Hoorick, Basile, et al.
Published: (2024) -
Differentiable Robot Rendering
by: Liu, Ruoshi, et al.
Published: (2024) -
EraseDraw: Learning to Draw Step-by-Step via Erasing Objects from Images
by: Canberk, Alper, et al.
Published: (2024)