Interacted Object Grounding in Spatio-Temporal Human-Object Interactions

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Xiaoyang, Wen, Boran, Liu, Xinpeng, Zhou, Zizheng, Fan, Hongwei, Lu, Cewu, Ma, Lizhuang, Chen, Yulong, Li, Yong-Lu
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913702952828928
author Liu, Xiaoyang
Wen, Boran
Liu, Xinpeng
Zhou, Zizheng
Fan, Hongwei
Lu, Cewu
Ma, Lizhuang
Chen, Yulong
Li, Yong-Lu
author_facet Liu, Xiaoyang
Wen, Boran
Liu, Xinpeng
Zhou, Zizheng
Fan, Hongwei
Lu, Cewu
Ma, Lizhuang
Chen, Yulong
Li, Yong-Lu
contents Spatio-temporal Human-Object Interaction (ST-HOI) understanding aims at detecting HOIs from videos, which is crucial for activity understanding. However, existing whole-body-object interaction video benchmarks overlook the truth that open-world objects are diverse, that is, they usually provide limited and predefined object classes. Therefore, we introduce a new open-world benchmark: Grounding Interacted Objects (GIO) including 1,098 interacted objects class and 290K interacted object boxes annotation. Accordingly, an object grounding task is proposed expecting vision systems to discover interacted objects. Even though today's detectors and grounding methods have succeeded greatly, they perform unsatisfactorily in localizing diverse and rare objects in GIO. This profoundly reveals the limitations of current vision systems and poses a great challenge. Thus, we explore leveraging spatio-temporal cues to address object grounding and propose a 4D question-answering framework (4D-QA) to discover interacted objects from diverse videos. Our method demonstrates significant superiority in extensive experiments compared to current baselines. Data and code will be publicly available at https://github.com/DirtyHarryLYL/HAKE-AVA.
format Preprint
id arxiv_https___arxiv_org_abs_2412_19542
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Interacted Object Grounding in Spatio-Temporal Human-Object Interactions
Liu, Xiaoyang
Wen, Boran
Liu, Xinpeng
Zhou, Zizheng
Fan, Hongwei
Lu, Cewu
Ma, Lizhuang
Chen, Yulong
Li, Yong-Lu
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Spatio-temporal Human-Object Interaction (ST-HOI) understanding aims at detecting HOIs from videos, which is crucial for activity understanding. However, existing whole-body-object interaction video benchmarks overlook the truth that open-world objects are diverse, that is, they usually provide limited and predefined object classes. Therefore, we introduce a new open-world benchmark: Grounding Interacted Objects (GIO) including 1,098 interacted objects class and 290K interacted object boxes annotation. Accordingly, an object grounding task is proposed expecting vision systems to discover interacted objects. Even though today's detectors and grounding methods have succeeded greatly, they perform unsatisfactorily in localizing diverse and rare objects in GIO. This profoundly reveals the limitations of current vision systems and poses a great challenge. Thus, we explore leveraging spatio-temporal cues to address object grounding and propose a 4D question-answering framework (4D-QA) to discover interacted objects from diverse videos. Our method demonstrates significant superiority in extensive experiments compared to current baselines. Data and code will be publicly available at https://github.com/DirtyHarryLYL/HAKE-AVA.
title Interacted Object Grounding in Spatio-Temporal Human-Object Interactions
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2412.19542