Synthesizing Visual Concepts as Vision-Language Programs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wüst, Antonia, Stammer, Wolfgang, Shindo, Hikaru, Helff, Lukas, Dhami, Devendra Singh, Kersting, Kristian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912726559752192
author Wüst, Antonia
Stammer, Wolfgang
Shindo, Hikaru
Helff, Lukas
Dhami, Devendra Singh
Kersting, Kristian
author_facet Wüst, Antonia
Stammer, Wolfgang
Shindo, Hikaru
Helff, Lukas
Dhami, Devendra Singh
Kersting, Kristian
contents Vision-Language models (VLMs) achieve strong performance on multimodal tasks but often fail at systematic visual reasoning tasks, leading to inconsistent or illogical outputs. Neuro-symbolic methods promise to address this by inducing interpretable logical rules, though they exploit rigid, domain-specific perception modules. We propose Vision-Language Programs (VLP), which combine the perceptual flexibility of VLMs with systematic reasoning of program synthesis. Rather than embedding reasoning inside the VLM, VLP leverages the model to produce structured visual descriptions that are compiled into neuro-symbolic programs. The resulting programs execute directly on images, remain consistent with task constraints, and provide human-interpretable explanations that enable easy shortcut mitigation. Experiments on synthetic and real-world datasets demonstrate that VLPs outperform direct and structured prompting, particularly on tasks requiring complex logical reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2511_18964
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Synthesizing Visual Concepts as Vision-Language Programs
Wüst, Antonia
Stammer, Wolfgang
Shindo, Hikaru
Helff, Lukas
Dhami, Devendra Singh
Kersting, Kristian
Artificial Intelligence
Vision-Language models (VLMs) achieve strong performance on multimodal tasks but often fail at systematic visual reasoning tasks, leading to inconsistent or illogical outputs. Neuro-symbolic methods promise to address this by inducing interpretable logical rules, though they exploit rigid, domain-specific perception modules. We propose Vision-Language Programs (VLP), which combine the perceptual flexibility of VLMs with systematic reasoning of program synthesis. Rather than embedding reasoning inside the VLM, VLP leverages the model to produce structured visual descriptions that are compiled into neuro-symbolic programs. The resulting programs execute directly on images, remain consistent with task constraints, and provide human-interpretable explanations that enable easy shortcut mitigation. Experiments on synthetic and real-world datasets demonstrate that VLPs outperform direct and structured prompting, particularly on tasks requiring complex logical reasoning.
title Synthesizing Visual Concepts as Vision-Language Programs
topic Artificial Intelligence
url https://arxiv.org/abs/2511.18964