Towards Unconstrained Human-Object Interaction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tonini, Francesco, Conti, Alessandro, Vaquero, Lorenzo, Beyan, Cigdem, Ricci, Elisa
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918448910565376
author Tonini, Francesco
Conti, Alessandro
Vaquero, Lorenzo
Beyan, Cigdem
Ricci, Elisa
author_facet Tonini, Francesco
Conti, Alessandro
Vaquero, Lorenzo
Beyan, Cigdem
Ricci, Elisa
contents Human-Object Interaction (HOI) detection is a longstanding computer vision problem concerned with predicting the interaction between humans and objects. Current HOI models rely on a vocabulary of interactions at training and inference time, limiting their applicability to static environments. With the advent of Multimodal Large Language Models (MLLMs), it has become feasible to explore more flexible paradigms for interaction recognition. In this work, we revisit HOI detection through the lens of MLLMs and apply them to in-the-wild HOI detection. We define the Unconstrained HOI (U-HOI) task, a novel HOI domain that removes the requirement for a predefined list of interactions at both training and inference. We evaluate a range of MLLMs on this setting and introduce a pipeline that includes test-time inference and language-to-graph conversion to extract structured interactions from free-form text. Our findings highlight the limitations of current HOI detectors and the value of MLLMs for U-HOI. Code will be available at https://github.com/francescotonini/anyhoi
format Preprint
id arxiv_https___arxiv_org_abs_2604_14069
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Towards Unconstrained Human-Object Interaction
Tonini, Francesco
Conti, Alessandro
Vaquero, Lorenzo
Beyan, Cigdem
Ricci, Elisa
Computer Vision and Pattern Recognition
Human-Object Interaction (HOI) detection is a longstanding computer vision problem concerned with predicting the interaction between humans and objects. Current HOI models rely on a vocabulary of interactions at training and inference time, limiting their applicability to static environments. With the advent of Multimodal Large Language Models (MLLMs), it has become feasible to explore more flexible paradigms for interaction recognition. In this work, we revisit HOI detection through the lens of MLLMs and apply them to in-the-wild HOI detection. We define the Unconstrained HOI (U-HOI) task, a novel HOI domain that removes the requirement for a predefined list of interactions at both training and inference. We evaluate a range of MLLMs on this setting and introduce a pipeline that includes test-time inference and language-to-graph conversion to extract structured interactions from free-form text. Our findings highlight the limitations of current HOI detectors and the value of MLLMs for U-HOI. Code will be available at https://github.com/francescotonini/anyhoi
title Towards Unconstrained Human-Object Interaction
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.14069