Foundation Model-Driven Framework for Human-Object Interaction Prediction with Segmentation Mask Integration

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Park, Juhan, Lee, Kyungjae, Chang, Hyung Jin, Cho, Jungchan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915263268519936
author Park, Juhan
Lee, Kyungjae
Chang, Hyung Jin
Cho, Jungchan
author_facet Park, Juhan
Lee, Kyungjae
Chang, Hyung Jin
Cho, Jungchan
contents In this work, we introduce Segmentation to Human-Object Interaction (\textit{\textbf{Seg2HOI}}) approach, a novel framework that integrates segmentation-based vision foundation models with the human-object interaction task, distinguished from traditional detection-based Human-Object Interaction (HOI) methods. Our approach enhances HOI detection by not only predicting the standard triplets but also introducing quadruplets, which extend HOI triplets by including segmentation masks for human-object pairs. More specifically, Seg2HOI inherits the properties of the vision foundation model (e.g., promptable and interactive mechanisms) and incorporates a decoder that applies these attributes to HOI task. Despite training only for HOI, without additional training mechanisms for these properties, the framework demonstrates that such features still operate efficiently. Extensive experiments on two public benchmark datasets demonstrate that Seg2HOI achieves performance comparable to state-of-the-art methods, even in zero-shot scenarios. Lastly, we propose that Seg2HOI can generate HOI quadruplets and interactive HOI segmentation from novel text and visual prompts that were not used during training, making it versatile for a wide range of applications by leveraging this flexibility.
format Preprint
id arxiv_https___arxiv_org_abs_2504_19847
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Foundation Model-Driven Framework for Human-Object Interaction Prediction with Segmentation Mask Integration
Park, Juhan
Lee, Kyungjae
Chang, Hyung Jin
Cho, Jungchan
Computer Vision and Pattern Recognition
Artificial Intelligence
In this work, we introduce Segmentation to Human-Object Interaction (\textit{\textbf{Seg2HOI}}) approach, a novel framework that integrates segmentation-based vision foundation models with the human-object interaction task, distinguished from traditional detection-based Human-Object Interaction (HOI) methods. Our approach enhances HOI detection by not only predicting the standard triplets but also introducing quadruplets, which extend HOI triplets by including segmentation masks for human-object pairs. More specifically, Seg2HOI inherits the properties of the vision foundation model (e.g., promptable and interactive mechanisms) and incorporates a decoder that applies these attributes to HOI task. Despite training only for HOI, without additional training mechanisms for these properties, the framework demonstrates that such features still operate efficiently. Extensive experiments on two public benchmark datasets demonstrate that Seg2HOI achieves performance comparable to state-of-the-art methods, even in zero-shot scenarios. Lastly, we propose that Seg2HOI can generate HOI quadruplets and interactive HOI segmentation from novel text and visual prompts that were not used during training, making it versatile for a wide range of applications by leveraging this flexibility.
title Foundation Model-Driven Framework for Human-Object Interaction Prediction with Segmentation Mask Integration
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2504.19847