InteractVLM: 3D Interaction Reasoning from 2D Foundational Models
Fuente:
arXiv
Saved in:
| Main Authors: | Dwivedi, Sai Kumar, Antić, Dimitrije, Tripathi, Shashank, Taheri, Omid, Schmid, Cordelia, Black, Michael J., Tzionas, Dimitrios |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SDFit: 3D Object Pose and Shape by Fitting a Morphable SDF to a Single Image
by: Antić, Dimitrije, et al.
Published: (2024)
by: Antić, Dimitrije, et al.
Published: (2024)
LEXIS: LatEnt ProXimal Interaction Signatures for 3D HOI from an Image
by: Antić, Dimitrije, et al.
Published: (2026)
by: Antić, Dimitrije, et al.
Published: (2026)
3D Whole-body Grasp Synthesis with Directional Controllability
by: Paschalidis, Georgios, et al.
Published: (2024)
by: Paschalidis, Georgios, et al.
Published: (2024)
PICO: Reconstructing 3D People In Contact with Objects
by: Cseke, Alpár, et al.
Published: (2025)
by: Cseke, Alpár, et al.
Published: (2025)
GRIP: Generating Interaction Poses Using Spatial Cues and Latent Consistency
by: Taheri, Omid, et al.
Published: (2023)
by: Taheri, Omid, et al.
Published: (2023)
Half-Physics: Enabling Kinematic 3D Human Model with Physical Interactions
by: Siyao, Li, et al.
Published: (2025)
by: Siyao, Li, et al.
Published: (2025)
InterDyn: Controllable Interactive Dynamics with Video Diffusion Models
by: Akkerman, Rick, et al.
Published: (2024)
by: Akkerman, Rick, et al.
Published: (2024)
HUMOS: Human Motion Model Conditioned on Body Shape
by: Tripathi, Shashank, et al.
Published: (2024)
by: Tripathi, Shashank, et al.
Published: (2024)
PuzzleAvatar: Assembling 3D Avatars from Personal Albums
by: Xiu, Yuliang, et al.
Published: (2024)
by: Xiu, Yuliang, et al.
Published: (2024)
CloSe: A 3D Clothing Segmentation Dataset and Model
by: Antić, Dimitrije, et al.
Published: (2024)
by: Antić, Dimitrije, et al.
Published: (2024)
PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction
by: Chen, Shizhe, et al.
Published: (2026)
by: Chen, Shizhe, et al.
Published: (2026)
ChatPose: Chatting about 3D Human Pose
by: Feng, Yao, et al.
Published: (2023)
by: Feng, Yao, et al.
Published: (2023)
A Versatile and Differentiable Hand-Object Interaction Representation
by: Morales, Théo, et al.
Published: (2024)
by: Morales, Théo, et al.
Published: (2024)
Predicting 4D Hand Trajectory from Monocular Videos
by: Ye, Yufei, et al.
Published: (2025)
by: Ye, Yufei, et al.
Published: (2025)
VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation
by: Bousselham, Walid, et al.
Published: (2025)
by: Bousselham, Walid, et al.
Published: (2025)
Towards Generalizable Vision-Language Robotic Manipulation: A Benchmark and LLM-guided 3D Policy
by: Garcia, Ricardo, et al.
Published: (2024)
by: Garcia, Ricardo, et al.
Published: (2024)
SUGAR: Pre-training 3D Visual Representations for Robotics
by: Chen, Shizhe, et al.
Published: (2024)
by: Chen, Shizhe, et al.
Published: (2024)
RHINO: Reconstructing Human Interactions with Novel Objects from Monocular Videos
by: Xue, Lixin, et al.
Published: (2026)
by: Xue, Lixin, et al.
Published: (2026)
Online 3D Scene Reconstruction Using Neural Object Priors
by: Chabal, Thomas, et al.
Published: (2025)
by: Chabal, Thomas, et al.
Published: (2025)
FirePlace: Geometric Refinements of LLM Common Sense Reasoning for 3D Object Placement
by: Huang, Ian, et al.
Published: (2025)
by: Huang, Ian, et al.
Published: (2025)
Reconstructing Objects along Hand Interaction Timelines in Egocentric Video
by: Zhu, Zhifan, et al.
Published: (2025)
by: Zhu, Zhifan, et al.
Published: (2025)
UniCon3R: Unified Contact-aware 4D Human-Scene Reconstruction from Monocular Video
by: Sur, Tanuj, et al.
Published: (2026)
by: Sur, Tanuj, et al.
Published: (2026)
FUSION: Full-Body Unified Motion Prior for Body and Hands via Diffusion
by: Duran, Enes, et al.
Published: (2026)
by: Duran, Enes, et al.
Published: (2026)
BrickNet: Graph-Backed Generative Brick Assembly
by: Kulits, Peter, et al.
Published: (2026)
by: Kulits, Peter, et al.
Published: (2026)
TokenHMR: Advancing Human Mesh Recovery with a Tokenized Pose Representation
by: Dwivedi, Sai Kumar, et al.
Published: (2024)
by: Dwivedi, Sai Kumar, et al.
Published: (2024)
Scaling Cross-Environment Failure Reasoning Data for Vision-Language Robotic Manipulation
by: Pacaud, Paul, et al.
Published: (2025)
by: Pacaud, Paul, et al.
Published: (2025)
HO-Flow: Generalizable Hand-Object Interaction Generation with Latent Flow Matching
by: Chen, Zerui, et al.
Published: (2026)
by: Chen, Zerui, et al.
Published: (2026)
UMO: Unified In-Context Learning Unlocks Motion Foundation Model Priors
by: Cong, Xiaoyan, et al.
Published: (2026)
by: Cong, Xiaoyan, et al.
Published: (2026)
Open-Vocabulary Functional 3D Human-Scene Interaction Generation
by: Liu, Jie, et al.
Published: (2026)
by: Liu, Jie, et al.
Published: (2026)
NIL: No-data Imitation Learning by Leveraging Pre-trained Video Diffusion Models
by: Albaba, Mert, et al.
Published: (2025)
by: Albaba, Mert, et al.
Published: (2025)
N3D-VLM: Native 3D Grounding Enables Accurate Spatial Reasoning in Vision-Language Models
by: Wang, Yuxin, et al.
Published: (2025)
by: Wang, Yuxin, et al.
Published: (2025)
Learning text-to-video retrieval from image captioning
by: Ventura, Lucas, et al.
Published: (2024)
by: Ventura, Lucas, et al.
Published: (2024)
Grounded Video Caption Generation
by: Kazakos, Evangelos, et al.
Published: (2024)
by: Kazakos, Evangelos, et al.
Published: (2024)
ComposeAnything: Composite Object Priors for Text-to-Image Generation
by: Khan, Zeeshan, et al.
Published: (2025)
by: Khan, Zeeshan, et al.
Published: (2025)
Large-scale Pre-training for Grounded Video Caption Generation
by: Kazakos, Evangelos, et al.
Published: (2025)
by: Kazakos, Evangelos, et al.
Published: (2025)
Retrieval-Enhanced Contrastive Vision-Text Models
by: Iscen, Ahmet, et al.
Published: (2023)
by: Iscen, Ahmet, et al.
Published: (2023)
OVFact: Measuring and Improving Open-Vocabulary Factuality for Long Caption Models
by: Wysoczańska, Monika, et al.
Published: (2025)
by: Wysoczańska, Monika, et al.
Published: (2025)
Moving by Looking: Towards Vision-Driven Avatar Motion Generation
by: Diomataris, Markos, et al.
Published: (2025)
by: Diomataris, Markos, et al.
Published: (2025)
OpenSU3D: Open World 3D Scene Understanding using Foundation Models
by: Mohiuddin, Rafay, et al.
Published: (2024)
by: Mohiuddin, Rafay, et al.
Published: (2024)
MoReVQA: Exploring Modular Reasoning Models for Video Question Answering
by: Min, Juhong, et al.
Published: (2024)
by: Min, Juhong, et al.
Published: (2024)
Similar Items
-
SDFit: 3D Object Pose and Shape by Fitting a Morphable SDF to a Single Image
by: Antić, Dimitrije, et al.
Published: (2024) -
LEXIS: LatEnt ProXimal Interaction Signatures for 3D HOI from an Image
by: Antić, Dimitrije, et al.
Published: (2026) -
3D Whole-body Grasp Synthesis with Directional Controllability
by: Paschalidis, Georgios, et al.
Published: (2024) -
PICO: Reconstructing 3D People In Contact with Objects
by: Cseke, Alpár, et al.
Published: (2025) -
GRIP: Generating Interaction Poses Using Spatial Cues and Latent Consistency
by: Taheri, Omid, et al.
Published: (2023)