Which objects help me to act effectively? Reasoning about physically-grounded affordances

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kemmeren, Anne, Burghouts, Gertjan, van Bekkum, Michael, Meijer, Wouter, van Mil, Jelle
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909262328889344
author Kemmeren, Anne
Burghouts, Gertjan
van Bekkum, Michael
Meijer, Wouter
van Mil, Jelle
author_facet Kemmeren, Anne
Burghouts, Gertjan
van Bekkum, Michael
Meijer, Wouter
van Mil, Jelle
contents For effective interactions with the open world, robots should understand how interactions with known and novel objects help them towards their goal. A key aspect of this understanding lies in detecting an object's affordances, which represent the potential effects that can be achieved by manipulating the object in various ways. Our approach leverages a dialogue of large language models (LLMs) and vision-language models (VLMs) to achieve open-world affordance detection. Given open-vocabulary descriptions of intended actions and effects, the useful objects in the environment are found. By grounding our system in the physical world, we account for the robot's embodiment and the intrinsic properties of the objects it encounters. In our experiments, we have shown that our method produces tailored outputs based on different embodiments or intended effects. The method was able to select a useful object from a set of distractors. Finetuning the VLM for physical properties improved overall performance. These results underline the importance of grounding the affordance search in the physical world, by taking into account robot embodiment and the physical properties of objects.
format Preprint
id arxiv_https___arxiv_org_abs_2407_13811
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Which objects help me to act effectively? Reasoning about physically-grounded affordances
Kemmeren, Anne
Burghouts, Gertjan
van Bekkum, Michael
Meijer, Wouter
van Mil, Jelle
Computer Vision and Pattern Recognition
Robotics
For effective interactions with the open world, robots should understand how interactions with known and novel objects help them towards their goal. A key aspect of this understanding lies in detecting an object's affordances, which represent the potential effects that can be achieved by manipulating the object in various ways. Our approach leverages a dialogue of large language models (LLMs) and vision-language models (VLMs) to achieve open-world affordance detection. Given open-vocabulary descriptions of intended actions and effects, the useful objects in the environment are found. By grounding our system in the physical world, we account for the robot's embodiment and the intrinsic properties of the objects it encounters. In our experiments, we have shown that our method produces tailored outputs based on different embodiments or intended effects. The method was able to select a useful object from a set of distractors. Finetuning the VLM for physical properties improved overall performance. These results underline the importance of grounding the affordance search in the physical world, by taking into account robot embodiment and the physical properties of objects.
title Which objects help me to act effectively? Reasoning about physically-grounded affordances
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2407.13811