QueryCraft: Transformer-Guided Query Initialization for Enhanced Human-Object Interaction Detection

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Yuxiao, Liang, Wolin, Lei, Yu, Xue, Weiying, Zhuang, Nan, Liu, Qi
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915441445699584
author Wang, Yuxiao
Liang, Wolin
Lei, Yu
Xue, Weiying
Zhuang, Nan
Liu, Qi
author_facet Wang, Yuxiao
Liang, Wolin
Lei, Yu
Xue, Weiying
Zhuang, Nan
Liu, Qi
contents Human-Object Interaction (HOI) detection aims to localize human-object pairs and recognize their interactions in images. Although DETR-based methods have recently emerged as the mainstream framework for HOI detection, they still suffer from a key limitation: Randomly initialized queries lack explicit semantics, leading to suboptimal detection performance. To address this challenge, we propose QueryCraft, a novel plug-and-play HOI detection framework that incorporates semantic priors and guided feature learning through transformer-based query initialization. Central to our approach is \textbf{ACTOR} (\textbf{A}ction-aware \textbf{C}ross-modal \textbf{T}ransf\textbf{OR}mer), a cross-modal Transformer encoder that jointly attends to visual regions and textual prompts to extract action-relevant features. Rather than merely aligning modalities, ACTOR leverages language-guided attention to infer interaction semantics and produce semantically meaningful query representations. To further enhance object-level query quality, we introduce a \textbf{P}erceptual \textbf{D}istilled \textbf{Q}uery \textbf{D}ecoder (\textbf{PDQD}), which distills object category awareness from a pre-trained detector to serve as object query initiation. This dual-branch query initialization enables the model to generate more interpretable and effective queries for HOI detection. Extensive experiments on HICO-Det and V-COCO benchmarks demonstrate that our method achieves state-of-the-art performance and strong generalization. Code will be released upon publication.
format Preprint
id arxiv_https___arxiv_org_abs_2508_08590
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle QueryCraft: Transformer-Guided Query Initialization for Enhanced Human-Object Interaction Detection
Wang, Yuxiao
Liang, Wolin
Lei, Yu
Xue, Weiying
Zhuang, Nan
Liu, Qi
Computer Vision and Pattern Recognition
Human-Computer Interaction
Human-Object Interaction (HOI) detection aims to localize human-object pairs and recognize their interactions in images. Although DETR-based methods have recently emerged as the mainstream framework for HOI detection, they still suffer from a key limitation: Randomly initialized queries lack explicit semantics, leading to suboptimal detection performance. To address this challenge, we propose QueryCraft, a novel plug-and-play HOI detection framework that incorporates semantic priors and guided feature learning through transformer-based query initialization. Central to our approach is \textbf{ACTOR} (\textbf{A}ction-aware \textbf{C}ross-modal \textbf{T}ransf\textbf{OR}mer), a cross-modal Transformer encoder that jointly attends to visual regions and textual prompts to extract action-relevant features. Rather than merely aligning modalities, ACTOR leverages language-guided attention to infer interaction semantics and produce semantically meaningful query representations. To further enhance object-level query quality, we introduce a \textbf{P}erceptual \textbf{D}istilled \textbf{Q}uery \textbf{D}ecoder (\textbf{PDQD}), which distills object category awareness from a pre-trained detector to serve as object query initiation. This dual-branch query initialization enables the model to generate more interpretable and effective queries for HOI detection. Extensive experiments on HICO-Det and V-COCO benchmarks demonstrate that our method achieves state-of-the-art performance and strong generalization. Code will be released upon publication.
title QueryCraft: Transformer-Guided Query Initialization for Enhanced Human-Object Interaction Detection
topic Computer Vision and Pattern Recognition
Human-Computer Interaction
url https://arxiv.org/abs/2508.08590