VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kang, Donggoo, Jeong, Dasol, Lee, Hyunmin, Park, Sangwoo, Park, Hasil, Kwon, Sunkyu, Kim, Yeongjoon, Paik, Joonki
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917849831833600
author Kang, Donggoo
Jeong, Dasol
Lee, Hyunmin
Park, Sangwoo
Park, Hasil
Kwon, Sunkyu
Kim, Yeongjoon
Paik, Joonki
author_facet Kang, Donggoo
Jeong, Dasol
Lee, Hyunmin
Park, Sangwoo
Park, Hasil
Kwon, Sunkyu
Kim, Yeongjoon
Paik, Joonki
contents The Large Vision Language Model (VLM) has recently addressed remarkable progress in bridging two fundamental modalities. VLM, trained by a sufficiently large dataset, exhibits a comprehensive understanding of both visual and linguistic to perform diverse tasks. To distill this knowledge accurately, in this paper, we introduce a novel approach that explicitly utilizes VLM as an objective function form for the Human-Object Interaction (HOI) detection task (\textbf{VLM-HOI}). Specifically, we propose a method that quantifies the similarity of the predicted HOI triplet using the Image-Text matching technique. We represent HOI triplets linguistically to fully utilize the language comprehension of VLMs, which are more suitable than CLIP models due to their localization and object-centric nature. This matching score is used as an objective for contrastive optimization. To our knowledge, this is the first utilization of VLM language abilities for HOI detection. Experiments demonstrate the effectiveness of our method, achieving state-of-the-art HOI detection accuracy on benchmarks. We believe integrating VLMs into HOI detection represents important progress towards more advanced and interpretable analysis of human-object interactions.
format Preprint
id arxiv_https___arxiv_org_abs_2411_18038
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis
Kang, Donggoo
Jeong, Dasol
Lee, Hyunmin
Park, Sangwoo
Park, Hasil
Kwon, Sunkyu
Kim, Yeongjoon
Paik, Joonki
Computer Vision and Pattern Recognition
Artificial Intelligence
The Large Vision Language Model (VLM) has recently addressed remarkable progress in bridging two fundamental modalities. VLM, trained by a sufficiently large dataset, exhibits a comprehensive understanding of both visual and linguistic to perform diverse tasks. To distill this knowledge accurately, in this paper, we introduce a novel approach that explicitly utilizes VLM as an objective function form for the Human-Object Interaction (HOI) detection task (\textbf{VLM-HOI}). Specifically, we propose a method that quantifies the similarity of the predicted HOI triplet using the Image-Text matching technique. We represent HOI triplets linguistically to fully utilize the language comprehension of VLMs, which are more suitable than CLIP models due to their localization and object-centric nature. This matching score is used as an objective for contrastive optimization. To our knowledge, this is the first utilization of VLM language abilities for HOI detection. Experiments demonstrate the effectiveness of our method, achieving state-of-the-art HOI detection accuracy on benchmarks. We believe integrating VLMs into HOI detection represents important progress towards more advanced and interpretable analysis of human-object interactions.
title VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2411.18038