Beyond Object Categories: Multi-Attribute Reference Understanding for Visual Grounding
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913756754214912 |
|---|---|
| author | Guo, Hao Zhu, Jianfei Fan, Wei Yi, Chunzhi Jiang, Feng |
| author_facet | Guo, Hao Zhu, Jianfei Fan, Wei Yi, Chunzhi Jiang, Feng |
| contents | Referring expression comprehension (REC) aims at achieving object localization based on natural language descriptions. However, existing REC approaches are constrained by object category descriptions and single-attribute intention descriptions, hindering their application in real-world scenarios. In natural human-robot interactions, users often express their desires through individual states and intentions, accompanied by guiding gestures, rather than detailed object descriptions. To address this challenge, we propose Multi-ref EC, a novel task framework that integrates state descriptions, derived intentions, and embodied gestures to locate target objects. We introduce the State-Intention-Gesture Attributes Reference (SIGAR) dataset, which combines state and intention expressions with embodied references. Through extensive experiments with various baseline models on SIGAR, we demonstrate that properly ordered multi-attribute references contribute to improved localization performance, revealing that single-attribute reference is insufficient for natural human-robot interaction scenarios. Our findings underscore the importance of multi-attribute reference expressions in advancing visual-language understanding. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_19240 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Beyond Object Categories: Multi-Attribute Reference Understanding for Visual Grounding Guo, Hao Zhu, Jianfei Fan, Wei Yi, Chunzhi Jiang, Feng Computer Vision and Pattern Recognition Human-Computer Interaction Referring expression comprehension (REC) aims at achieving object localization based on natural language descriptions. However, existing REC approaches are constrained by object category descriptions and single-attribute intention descriptions, hindering their application in real-world scenarios. In natural human-robot interactions, users often express their desires through individual states and intentions, accompanied by guiding gestures, rather than detailed object descriptions. To address this challenge, we propose Multi-ref EC, a novel task framework that integrates state descriptions, derived intentions, and embodied gestures to locate target objects. We introduce the State-Intention-Gesture Attributes Reference (SIGAR) dataset, which combines state and intention expressions with embodied references. Through extensive experiments with various baseline models on SIGAR, we demonstrate that properly ordered multi-attribute references contribute to improved localization performance, revealing that single-attribute reference is insufficient for natural human-robot interaction scenarios. Our findings underscore the importance of multi-attribute reference expressions in advancing visual-language understanding. |
| title | Beyond Object Categories: Multi-Attribute Reference Understanding for Visual Grounding |
| topic | Computer Vision and Pattern Recognition Human-Computer Interaction |
| url | https://arxiv.org/abs/2503.19240 |