Bongards at the Boundary of Perception and Reasoning: Programs or Language?

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Langenfeld, Cassidy, Beger, Claas, Geng, Gloria, Piriyakulkij, Wasu Top, Hu, Keya, Pu, Yewen, Ellis, Kevin
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912871636533248
author Langenfeld, Cassidy
Beger, Claas
Geng, Gloria
Piriyakulkij, Wasu Top
Hu, Keya
Pu, Yewen
Ellis, Kevin
author_facet Langenfeld, Cassidy
Beger, Claas
Geng, Gloria
Piriyakulkij, Wasu Top
Hu, Keya
Pu, Yewen
Ellis, Kevin
contents Vision-Language Models (VLMs) have made great strides in everyday visual tasks, such as captioning a natural image, or answering commonsense questions about such images. But humans possess the puzzling ability to deploy their visual reasoning abilities in radically new situations, a skill rigorously tested by the classic set of visual reasoning challenges known as the Bongard problems. We present a neurosymbolic approach to solving these problems: given a hypothesized solution rule for a Bongard problem, we leverage LLMs to generate parameterized programmatic representations for the rule and perform parameter fitting using Bayesian optimization. We evaluate our method on classifying Bongard problem images given the ground truth rule, as well as on solving the problems from scratch.
format Preprint
id arxiv_https___arxiv_org_abs_2602_03038
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Bongards at the Boundary of Perception and Reasoning: Programs or Language?
Langenfeld, Cassidy
Beger, Claas
Geng, Gloria
Piriyakulkij, Wasu Top
Hu, Keya
Pu, Yewen
Ellis, Kevin
Computer Vision and Pattern Recognition
Artificial Intelligence
Vision-Language Models (VLMs) have made great strides in everyday visual tasks, such as captioning a natural image, or answering commonsense questions about such images. But humans possess the puzzling ability to deploy their visual reasoning abilities in radically new situations, a skill rigorously tested by the classic set of visual reasoning challenges known as the Bongard problems. We present a neurosymbolic approach to solving these problems: given a hypothesized solution rule for a Bongard problem, we leverage LLMs to generate parameterized programmatic representations for the rule and perform parameter fitting using Bayesian optimization. We evaluate our method on classifying Bongard problem images given the ground truth rule, as well as on solving the problems from scratch.
title Bongards at the Boundary of Perception and Reasoning: Programs or Language?
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2602.03038