VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis
Fuente:
arXiv
Saved in:
| Main Authors: | Kang, Donggoo, Jeong, Dasol, Lee, Hyunmin, Park, Sangwoo, Park, Hasil, Kwon, Sunkyu, Kim, Yeongjoon, Paik, Joonki |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Structure-Preserving Zero-Shot Image Editing via Stage-Wise Latent Injection in Diffusion Models
by: Jeong, Dasol, et al.
Published: (2025)
by: Jeong, Dasol, et al.
Published: (2025)
Consistent Zero-shot 3D Texture Synthesis Using Geometry-aware Diffusion and Temporal Video Models
by: Kang, Donggoo, et al.
Published: (2025)
by: Kang, Donggoo, et al.
Published: (2025)
LEAP:D -- A Novel Prompt-based Approach for Domain-Generalized Aerial Object Detection
by: Park, Chanyeong, et al.
Published: (2024)
by: Park, Chanyeong, et al.
Published: (2024)
CL-HOI: Cross-Level Human-Object Interaction Distillation from Vision Large Language Models
by: Gao, Jianjun, et al.
Published: (2024)
by: Gao, Jianjun, et al.
Published: (2024)
HOI-Ref: Hand-Object Interaction Referral in Egocentric Vision
by: Bansal, Siddhant, et al.
Published: (2024)
by: Bansal, Siddhant, et al.
Published: (2024)
Ins-HOI: Instance Aware Human-Object Interactions Recovery
by: Zhang, Jiajun, et al.
Published: (2023)
by: Zhang, Jiajun, et al.
Published: (2023)
ViHOI: Human-Object Interaction Synthesis with Visual Priors
by: Cai, Songjin, et al.
Published: (2026)
by: Cai, Songjin, et al.
Published: (2026)
OneHOI: Unifying Human-Object Interaction Generation and Editing
by: Hoe, Jiun Tian, et al.
Published: (2026)
by: Hoe, Jiun Tian, et al.
Published: (2026)
HOI4ABOT: Human-Object Interaction Anticipation for Human Intention Reading Collaborative roBOTs
by: Mascaro, Esteve Valls, et al.
Published: (2023)
by: Mascaro, Esteve Valls, et al.
Published: (2023)
PA-HOI: A Physics-Aware Human and Object Interaction Dataset
by: Wang, Ruiyan, et al.
Published: (2025)
by: Wang, Ruiyan, et al.
Published: (2025)
ContextHOI: Spatial Context Learning for Human-Object Interaction Detection
by: Jia, Mingda, et al.
Published: (2024)
by: Jia, Mingda, et al.
Published: (2024)
HOI-Dyn: Learning Interaction Dynamics for Human-Object Motion Diffusion
by: Wu, Lin, et al.
Published: (2025)
by: Wu, Lin, et al.
Published: (2025)
RoHOI: Robustness Benchmark for Human-Object Interaction Detection
by: Wen, Di, et al.
Published: (2025)
by: Wen, Di, et al.
Published: (2025)
OnlineHOI: Towards Online Human-Object Interaction Generation and Perception
by: Ji, Yihong, et al.
Published: (2025)
by: Ji, Yihong, et al.
Published: (2025)
PersonaHOI: Effortlessly Improving Personalized Face with Human-Object Interaction Generation
by: Hu, Xinting, et al.
Published: (2025)
by: Hu, Xinting, et al.
Published: (2025)
EZ-HOI: VLM Adaptation via Guided Prompt Learning for Zero-Shot HOI Detection
by: Lei, Qinqian, et al.
Published: (2024)
by: Lei, Qinqian, et al.
Published: (2024)
RT-VLM: Re-Thinking Vision Language Model with 4-Clues for Real-World Object Recognition Robustness
by: Park, Junghyun, et al.
Published: (2025)
by: Park, Junghyun, et al.
Published: (2025)
TeamHOI: Learning a Unified Policy for Cooperative Human-Object Interactions with Any Team Size
by: Lionar, Stefan, et al.
Published: (2026)
by: Lionar, Stefan, et al.
Published: (2026)
VLM-SubtleBench: How Far Are VLMs from Human-Level Subtle Comparative Reasoning?
by: Kim, Minkyu, et al.
Published: (2026)
by: Kim, Minkyu, et al.
Published: (2026)
HOI-R1: Exploring the Potential of Multimodal Large Language Models for Human-Object Interaction Detection
by: Chen, Junwen, et al.
Published: (2025)
by: Chen, Junwen, et al.
Published: (2025)
HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction
by: Bao, Chen, et al.
Published: (2024)
by: Bao, Chen, et al.
Published: (2024)
HOI-M3:Capture Multiple Humans and Objects Interaction within Contextual Environment
by: Zhang, Juze, et al.
Published: (2024)
by: Zhang, Juze, et al.
Published: (2024)
CycleHOI: Improving Human-Object Interaction Detection with Cycle Consistency of Detection and Generation
by: Wang, Yisen, et al.
Published: (2024)
by: Wang, Yisen, et al.
Published: (2024)
ChainHOI: Joint-based Kinematic Chain Modeling for Human-Object Interaction Generation
by: Zeng, Ling-An, et al.
Published: (2025)
by: Zeng, Ling-An, et al.
Published: (2025)
Evaluating Visual and Cultural Interpretation: The K-Viscuit Benchmark with Human-VLM Collaboration
by: Park, ChaeHun, et al.
Published: (2024)
by: Park, ChaeHun, et al.
Published: (2024)
HOI-Swap: Swapping Objects in Videos with Hand-Object Interaction Awareness
by: Xue, Zihui, et al.
Published: (2024)
by: Xue, Zihui, et al.
Published: (2024)
HOI-PAGE: Zero-Shot Human-Object Interaction Generation with Part Affordance Guidance
by: Li, Lei, et al.
Published: (2025)
by: Li, Lei, et al.
Published: (2025)
ROK Defense M&S in the Age of Hyperscale AI: Concepts, Challenges, and Future Directions
by: Lee, Youngjoon, et al.
Published: (2024)
by: Lee, Youngjoon, et al.
Published: (2024)
EgoWorld: Translating Exocentric View to Egocentric View using Rich Exocentric Observations
by: Park, Junho, et al.
Published: (2025)
by: Park, Junho, et al.
Published: (2025)
CrossHOI-Bench: A Unified Benchmark for HOI Evaluation across Vision-Language Models and HOI-Specific Methods
by: Lei, Qinqian, et al.
Published: (2025)
by: Lei, Qinqian, et al.
Published: (2025)
GenHOI: Generalizing Text-driven 4D Human-Object Interaction Synthesis for Unseen Objects
by: Li, Shujia, et al.
Published: (2025)
by: Li, Shujia, et al.
Published: (2025)
UniHOI: Unified Human-Object Interaction Understanding via Unified Token Space
by: Yang, Panqi, et al.
Published: (2025)
by: Yang, Panqi, et al.
Published: (2025)
Topological Alignment of Shared Vision-Language Embedding Space
by: You, Junwon, et al.
Published: (2025)
by: You, Junwon, et al.
Published: (2025)
I'M HOI: Inertia-aware Monocular Capture of 3D Human-Object Interactions
by: Zhao, Chengfeng, et al.
Published: (2023)
by: Zhao, Chengfeng, et al.
Published: (2023)
ScoreHOI: Physically Plausible Reconstruction of Human-Object Interaction via Score-Guided Diffusion
by: Li, Ao, et al.
Published: (2025)
by: Li, Ao, et al.
Published: (2025)
Uni-HOI:A Unified framework for Learning the Joint distribution of Text and Human-Object Interaction
by: Zhang, Mengfei, et al.
Published: (2026)
by: Zhang, Mengfei, et al.
Published: (2026)
ScriptHOI: Learning Scripted State Transitions for Open-Vocabulary Human-Object Interaction Detection
by: Nguyen, Minh Anh, et al.
Published: (2026)
by: Nguyen, Minh Anh, et al.
Published: (2026)
DreamHOI: Subject-Driven Generation of 3D Human-Object Interactions with Diffusion Priors
by: Zhu, Thomas Hanwen, et al.
Published: (2024)
by: Zhu, Thomas Hanwen, et al.
Published: (2024)
F-HOI: Toward Fine-grained Semantic-Aligned 3D Human-Object Interactions
by: Yang, Jie, et al.
Published: (2024)
by: Yang, Jie, et al.
Published: (2024)
Read Like a Radiologist: Efficient Vision-Language Model for 3D Medical Imaging Interpretation
by: Lee, Changsun, et al.
Published: (2024)
by: Lee, Changsun, et al.
Published: (2024)
Similar Items
-
Structure-Preserving Zero-Shot Image Editing via Stage-Wise Latent Injection in Diffusion Models
by: Jeong, Dasol, et al.
Published: (2025) -
Consistent Zero-shot 3D Texture Synthesis Using Geometry-aware Diffusion and Temporal Video Models
by: Kang, Donggoo, et al.
Published: (2025) -
LEAP:D -- A Novel Prompt-based Approach for Domain-Generalized Aerial Object Detection
by: Park, Chanyeong, et al.
Published: (2024) -
CL-HOI: Cross-Level Human-Object Interaction Distillation from Vision Large Language Models
by: Gao, Jianjun, et al.
Published: (2024) -
HOI-Ref: Hand-Object Interaction Referral in Egocentric Vision
by: Bansal, Siddhant, et al.
Published: (2024)