Human-Object Interaction via Automatically Designed VLM-Guided Motion Policy
Fuente:
arXiv
Saved in:
| Main Authors: | Deng, Zekai, Shi, Ye, Ji, Kaiyang, Xu, Lan, Huang, Shaoli, Wang, Jingya |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Unified Diffusion Framework for Scene-aware Human Motion Estimation from Sparse Signals
by: Tang, Jiangnan, et al.
Published: (2024)
by: Tang, Jiangnan, et al.
Published: (2024)
Towards Immersive Human-X Interaction: A Real-Time Framework for Physically Plausible Motion Synthesis
by: Ji, Kaiyang, et al.
Published: (2025)
by: Ji, Kaiyang, et al.
Published: (2025)
THOR: Text to Human-Object Interaction Diffusion via Relation Intervention
by: Wu, Qianyang, et al.
Published: (2024)
by: Wu, Qianyang, et al.
Published: (2024)
Guiding Human-Object Interactions with Rich Geometry and Relations
by: Xue, Mengqing, et al.
Published: (2025)
by: Xue, Mengqing, et al.
Published: (2025)
ARFlow: Human Action-Reaction Flow Matching with Physical Guidance
by: Jiang, Wentao, et al.
Published: (2025)
by: Jiang, Wentao, et al.
Published: (2025)
Monocular Human-Object Reconstruction in the Wild
by: Huo, Chaofan, et al.
Published: (2024)
by: Huo, Chaofan, et al.
Published: (2024)
HOI-M3:Capture Multiple Humans and Objects Interaction within Contextual Environment
by: Zhang, Juze, et al.
Published: (2024)
by: Zhang, Juze, et al.
Published: (2024)
Gaze-guided Hand-Object Interaction Synthesis: Dataset and Method
by: Tian, Jie, et al.
Published: (2024)
by: Tian, Jie, et al.
Published: (2024)
StackFLOW: Monocular Human-Object Reconstruction by Stacked Normalizing Flow with Offset
by: Huo, Chaofan, et al.
Published: (2024)
by: Huo, Chaofan, et al.
Published: (2024)
Realistic Human Motion Generation with Cross-Diffusion Models
by: Ren, Zeping, et al.
Published: (2023)
by: Ren, Zeping, et al.
Published: (2023)
Bilateral Collaboration with Large Vision-Language Models for Open Vocabulary Human-Object Interaction Detection
by: Hu, Yupeng, et al.
Published: (2025)
by: Hu, Yupeng, et al.
Published: (2025)
DiscoForcing: A Unified Framework for Real-Time Audio-Driven Character Control with Diffusion Forcing
by: Ji, Kaiyang, et al.
Published: (2026)
by: Ji, Kaiyang, et al.
Published: (2026)
I'M HOI: Inertia-aware Monocular Capture of 3D Human-Object Interactions
by: Zhao, Chengfeng, et al.
Published: (2023)
by: Zhao, Chengfeng, et al.
Published: (2023)
HoloGest: Decoupled Diffusion and Motion Priors for Generating Holisticly Expressive Co-speech Gestures
by: Cheng, Yongkang, et al.
Published: (2025)
by: Cheng, Yongkang, et al.
Published: (2025)
OpenHOI: Open-World Hand-Object Interaction Synthesis with Multimodal Large Language Model
by: Zhang, Zhenhao, et al.
Published: (2025)
by: Zhang, Zhenhao, et al.
Published: (2025)
HOI-Dyn: Learning Interaction Dynamics for Human-Object Motion Diffusion
by: Wu, Lin, et al.
Published: (2025)
by: Wu, Lin, et al.
Published: (2025)
Stability-Driven Motion Generation for Object-Guided Human-Human Co-Manipulation
by: Xu, Jiahao, et al.
Published: (2026)
by: Xu, Jiahao, et al.
Published: (2026)
OmniMotion-X: Versatile Multimodal Whole-Body Motion Generation
by: Xu, Guowei, et al.
Published: (2025)
by: Xu, Guowei, et al.
Published: (2025)
Efficient and Scalable Monocular Human-Object Interaction Motion Reconstruction
by: Wen, Boran, et al.
Published: (2025)
by: Wen, Boran, et al.
Published: (2025)
RopeTP: Global Human Motion Recovery via Integrating Robust Pose Estimation with Diffusion Trajectory Prior
by: Liang, Mingjiang, et al.
Published: (2024)
by: Liang, Mingjiang, et al.
Published: (2024)
Programmable Motion Generation for Open-Set Motion Control Tasks
by: Liu, Hanchao, et al.
Published: (2024)
by: Liu, Hanchao, et al.
Published: (2024)
Unsupervised Cross-Domain Image Retrieval via Prototypical Optimal Transport
by: Li, Bin, et al.
Published: (2024)
by: Li, Bin, et al.
Published: (2024)
DISPLAY: Directable Human-Object Interaction Video Generation via Sparse Motion Guidance and Multi-Task Auxiliary
by: Guan, Jiazhi, et al.
Published: (2026)
by: Guan, Jiazhi, et al.
Published: (2026)
Contextually Affinitive Neighborhood Refinery for Deep Clustering
by: Yu, Chunlin, et al.
Published: (2023)
by: Yu, Chunlin, et al.
Published: (2023)
Interaction Replica: Tracking Human-Object Interaction and Scene Changes From Human Motion
by: Guzov, Vladimir, et al.
Published: (2022)
by: Guzov, Vladimir, et al.
Published: (2022)
SignAvatars: A Large-scale 3D Sign Language Holistic Motion Dataset and Benchmark
by: Yu, Zhengdi, et al.
Published: (2023)
by: Yu, Zhengdi, et al.
Published: (2023)
HandDiffuse: Generative Controllers for Two-Hand Interactions via Diffusion Models
by: Lin, Pei, et al.
Published: (2023)
by: Lin, Pei, et al.
Published: (2023)
InteractAnything: Zero-shot Human Object Interaction Synthesis via LLM Feedback and Object Affordance Parsing
by: Zhang, Jinlu, et al.
Published: (2025)
by: Zhang, Jinlu, et al.
Published: (2025)
SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents
by: Huang-Menders, Alexander, et al.
Published: (2025)
by: Huang-Menders, Alexander, et al.
Published: (2025)
VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis
by: Kang, Donggoo, et al.
Published: (2024)
by: Kang, Donggoo, et al.
Published: (2024)
FlashCap: Millisecond-Accurate Human Motion Capture via Flashing LEDs and Event-Based Vision
by: Wu, Zekai, et al.
Published: (2026)
by: Wu, Zekai, et al.
Published: (2026)
HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction
by: Bao, Chen, et al.
Published: (2024)
by: Bao, Chen, et al.
Published: (2024)
Multimodal Graph Network Modeling for Human-Object Interaction Detection with PDE Graph Diffusion
by: Ji, Wenxuan, et al.
Published: (2025)
by: Ji, Wenxuan, et al.
Published: (2025)
Towards Motion Turing Test: Evaluating Human-Likeness in Humanoid Robots
by: Li, Mingzhe, et al.
Published: (2026)
by: Li, Mingzhe, et al.
Published: (2026)
CAST: Mitigating Object Hallucination in Large Vision-Language Models via Caption-Guided Visual Attention Steering
by: Li, Qiming, et al.
Published: (2026)
by: Li, Qiming, et al.
Published: (2026)
Text-Guided 6D Object Pose Rearrangement via Closed-Loop VLM Agents
by: Baik, Sangwon, et al.
Published: (2026)
by: Baik, Sangwon, et al.
Published: (2026)
SMGDiff: Soccer Motion Generation using diffusion probabilistic models
by: Yang, Hongdi, et al.
Published: (2024)
by: Yang, Hongdi, et al.
Published: (2024)
EchoMotion: Unified Human Video and Motion Generation via Dual-Modality Diffusion Transformer
by: Yang, Yuxiao, et al.
Published: (2025)
by: Yang, Yuxiao, et al.
Published: (2025)
Egocentric Human-Object Interaction Detection: A New Benchmark and Method
by: Deng, Kunyuan, et al.
Published: (2025)
by: Deng, Kunyuan, et al.
Published: (2025)
HOID-R1: Reinforcement Learning for Open-World Human-Object Interaction Detection Reasoning with Multimodal Large Language Model
by: Zhang, Zhenhao, et al.
Published: (2025)
by: Zhang, Zhenhao, et al.
Published: (2025)
Similar Items
-
A Unified Diffusion Framework for Scene-aware Human Motion Estimation from Sparse Signals
by: Tang, Jiangnan, et al.
Published: (2024) -
Towards Immersive Human-X Interaction: A Real-Time Framework for Physically Plausible Motion Synthesis
by: Ji, Kaiyang, et al.
Published: (2025) -
THOR: Text to Human-Object Interaction Diffusion via Relation Intervention
by: Wu, Qianyang, et al.
Published: (2024) -
Guiding Human-Object Interactions with Rich Geometry and Relations
by: Xue, Mengqing, et al.
Published: (2025) -
ARFlow: Human Action-Reaction Flow Matching with Physical Guidance
by: Jiang, Wentao, et al.
Published: (2025)