PropTest: Automatic Property Testing for Improved Visual Programming

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Koo, Jaywon, Yang, Ziyan, Cascante-Bonilla, Paola, Ray, Baishakhi, Ordonez, Vicente
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910538646159360
author Koo, Jaywon
Yang, Ziyan
Cascante-Bonilla, Paola
Ray, Baishakhi
Ordonez, Vicente
author_facet Koo, Jaywon
Yang, Ziyan
Cascante-Bonilla, Paola
Ray, Baishakhi
Ordonez, Vicente
contents Visual Programming has recently emerged as an alternative to end-to-end black-box visual reasoning models. This type of method leverages Large Language Models (LLMs) to generate the source code for an executable computer program that solves a given problem. This strategy has the advantage of offering an interpretable reasoning path and does not require finetuning a model with task-specific data. We propose PropTest, a general strategy that improves visual programming by further using an LLM to generate code that tests for visual properties in an initial round of proposed solutions. Our method generates tests for data-type consistency, output syntax, and semantic properties. PropTest achieves comparable results to state-of-the-art methods while using publicly available LLMs. This is demonstrated across different benchmarks on visual question answering and referring expression comprehension. Particularly, PropTest improves ViperGPT by obtaining 46.1\% accuracy (+6.0\%) on GQA using Llama3-8B and 59.5\% (+8.1\%) on RefCOCO+ using CodeLlama-34B.
format Preprint
id arxiv_https___arxiv_org_abs_2403_16921
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle PropTest: Automatic Property Testing for Improved Visual Programming
Koo, Jaywon
Yang, Ziyan
Cascante-Bonilla, Paola
Ray, Baishakhi
Ordonez, Vicente
Computer Vision and Pattern Recognition
Visual Programming has recently emerged as an alternative to end-to-end black-box visual reasoning models. This type of method leverages Large Language Models (LLMs) to generate the source code for an executable computer program that solves a given problem. This strategy has the advantage of offering an interpretable reasoning path and does not require finetuning a model with task-specific data. We propose PropTest, a general strategy that improves visual programming by further using an LLM to generate code that tests for visual properties in an initial round of proposed solutions. Our method generates tests for data-type consistency, output syntax, and semantic properties. PropTest achieves comparable results to state-of-the-art methods while using publicly available LLMs. This is demonstrated across different benchmarks on visual question answering and referring expression comprehension. Particularly, PropTest improves ViperGPT by obtaining 46.1\% accuracy (+6.0\%) on GQA using Llama3-8B and 59.5\% (+8.1\%) on RefCOCO+ using CodeLlama-34B.
title PropTest: Automatic Property Testing for Improved Visual Programming
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.16921