TEXT2AFFORD: Probing Object Affordance Prediction abilities of Language Models solely from Text

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Adak, Sayantan, Agrawal, Daivik, Mukherjee, Animesh, Aditya, Somak
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912606597414912
author Adak, Sayantan
Agrawal, Daivik
Mukherjee, Animesh
Aditya, Somak
author_facet Adak, Sayantan
Agrawal, Daivik
Mukherjee, Animesh
Aditya, Somak
contents We investigate the knowledge of object affordances in pre-trained language models (LMs) and pre-trained Vision-Language models (VLMs). A growing body of literature shows that PTLMs fail inconsistently and non-intuitively, demonstrating a lack of reasoning and grounding. To take a first step toward quantifying the effect of grounding (or lack thereof), we curate a novel and comprehensive dataset of object affordances -- Text2Afford, characterized by 15 affordance classes. Unlike affordance datasets collected in vision and language domains, we annotate in-the-wild sentences with objects and affordances. Experimental results reveal that PTLMs exhibit limited reasoning abilities when it comes to uncommon object affordances. We also observe that pre-trained VLMs do not necessarily capture object affordances effectively. Through few-shot fine-tuning, we demonstrate improvement in affordance knowledge in PTLMs and VLMs. Our research contributes a novel dataset for language grounding tasks, and presents insights into LM capabilities, advancing the understanding of object affordances. Codes and data are available at https://github.com/sayantan11995/Text2Afford
format Preprint
id arxiv_https___arxiv_org_abs_2402_12881
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle TEXT2AFFORD: Probing Object Affordance Prediction abilities of Language Models solely from Text
Adak, Sayantan
Agrawal, Daivik
Mukherjee, Animesh
Aditya, Somak
Computation and Language
We investigate the knowledge of object affordances in pre-trained language models (LMs) and pre-trained Vision-Language models (VLMs). A growing body of literature shows that PTLMs fail inconsistently and non-intuitively, demonstrating a lack of reasoning and grounding. To take a first step toward quantifying the effect of grounding (or lack thereof), we curate a novel and comprehensive dataset of object affordances -- Text2Afford, characterized by 15 affordance classes. Unlike affordance datasets collected in vision and language domains, we annotate in-the-wild sentences with objects and affordances. Experimental results reveal that PTLMs exhibit limited reasoning abilities when it comes to uncommon object affordances. We also observe that pre-trained VLMs do not necessarily capture object affordances effectively. Through few-shot fine-tuning, we demonstrate improvement in affordance knowledge in PTLMs and VLMs. Our research contributes a novel dataset for language grounding tasks, and presents insights into LM capabilities, advancing the understanding of object affordances. Codes and data are available at https://github.com/sayantan11995/Text2Afford
title TEXT2AFFORD: Probing Object Affordance Prediction abilities of Language Models solely from Text
topic Computation and Language
url https://arxiv.org/abs/2402.12881