Beyond Object Categories: Multi-Attribute Reference Understanding for Visual Grounding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guo, Hao, Zhu, Jianfei, Fan, Wei, Yi, Chunzhi, Jiang, Feng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913756754214912
author Guo, Hao
Zhu, Jianfei
Fan, Wei
Yi, Chunzhi
Jiang, Feng
author_facet Guo, Hao
Zhu, Jianfei
Fan, Wei
Yi, Chunzhi
Jiang, Feng
contents Referring expression comprehension (REC) aims at achieving object localization based on natural language descriptions. However, existing REC approaches are constrained by object category descriptions and single-attribute intention descriptions, hindering their application in real-world scenarios. In natural human-robot interactions, users often express their desires through individual states and intentions, accompanied by guiding gestures, rather than detailed object descriptions. To address this challenge, we propose Multi-ref EC, a novel task framework that integrates state descriptions, derived intentions, and embodied gestures to locate target objects. We introduce the State-Intention-Gesture Attributes Reference (SIGAR) dataset, which combines state and intention expressions with embodied references. Through extensive experiments with various baseline models on SIGAR, we demonstrate that properly ordered multi-attribute references contribute to improved localization performance, revealing that single-attribute reference is insufficient for natural human-robot interaction scenarios. Our findings underscore the importance of multi-attribute reference expressions in advancing visual-language understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2503_19240
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond Object Categories: Multi-Attribute Reference Understanding for Visual Grounding
Guo, Hao
Zhu, Jianfei
Fan, Wei
Yi, Chunzhi
Jiang, Feng
Computer Vision and Pattern Recognition
Human-Computer Interaction
Referring expression comprehension (REC) aims at achieving object localization based on natural language descriptions. However, existing REC approaches are constrained by object category descriptions and single-attribute intention descriptions, hindering their application in real-world scenarios. In natural human-robot interactions, users often express their desires through individual states and intentions, accompanied by guiding gestures, rather than detailed object descriptions. To address this challenge, we propose Multi-ref EC, a novel task framework that integrates state descriptions, derived intentions, and embodied gestures to locate target objects. We introduce the State-Intention-Gesture Attributes Reference (SIGAR) dataset, which combines state and intention expressions with embodied references. Through extensive experiments with various baseline models on SIGAR, we demonstrate that properly ordered multi-attribute references contribute to improved localization performance, revealing that single-attribute reference is insufficient for natural human-robot interaction scenarios. Our findings underscore the importance of multi-attribute reference expressions in advancing visual-language understanding.
title Beyond Object Categories: Multi-Attribute Reference Understanding for Visual Grounding
topic Computer Vision and Pattern Recognition
Human-Computer Interaction
url https://arxiv.org/abs/2503.19240