Open-World Visual Reasoning by a Neuro-Symbolic Program of Zero-Shot Symbols

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Burghouts, Gertjan, Hillerström, Fieke, Walraven, Erwin, van Bekkum, Michael, Ruis, Frank, Sijs, Joris, van Mil, Jelle, Dijk, Judith
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913435246133248
author Burghouts, Gertjan
Hillerström, Fieke
Walraven, Erwin
van Bekkum, Michael
Ruis, Frank
Sijs, Joris
van Mil, Jelle
Dijk, Judith
author_facet Burghouts, Gertjan
Hillerström, Fieke
Walraven, Erwin
van Bekkum, Michael
Ruis, Frank
Sijs, Joris
van Mil, Jelle
Dijk, Judith
contents We consider the problem of finding spatial configurations of multiple objects in images, e.g., a mobile inspection robot is tasked to localize abandoned tools on the floor. We define the spatial configuration of objects by first-order logic in terms of relations and attributes. A neuro-symbolic program matches the logic formulas to probabilistic object proposals for the given image, provided by language-vision models by querying them for the symbols. This work is the first to combine neuro-symbolic programming (reasoning) and language-vision models (learning) to find spatial configurations of objects in images in an open world setting. We show the effectiveness by finding abandoned tools on floors and leaking pipes. We find that most prediction errors are due to biases in the language-vision model.
format Preprint
id arxiv_https___arxiv_org_abs_2407_13382
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Open-World Visual Reasoning by a Neuro-Symbolic Program of Zero-Shot Symbols
Burghouts, Gertjan
Hillerström, Fieke
Walraven, Erwin
van Bekkum, Michael
Ruis, Frank
Sijs, Joris
van Mil, Jelle
Dijk, Judith
Machine Learning
We consider the problem of finding spatial configurations of multiple objects in images, e.g., a mobile inspection robot is tasked to localize abandoned tools on the floor. We define the spatial configuration of objects by first-order logic in terms of relations and attributes. A neuro-symbolic program matches the logic formulas to probabilistic object proposals for the given image, provided by language-vision models by querying them for the symbols. This work is the first to combine neuro-symbolic programming (reasoning) and language-vision models (learning) to find spatial configurations of objects in images in an open world setting. We show the effectiveness by finding abandoned tools on floors and leaking pipes. We find that most prediction errors are due to biases in the language-vision model.
title Open-World Visual Reasoning by a Neuro-Symbolic Program of Zero-Shot Symbols
topic Machine Learning
url https://arxiv.org/abs/2407.13382