De-fine: Decomposing and Refining Visual Programs with Auto-Feedback

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gao, Minghe, Li, Juncheng, Fei, Hao, Pang, Liang, Ji, Wei, Wang, Guoming, Lv, Zheqi, Zhang, Wenqiao, Tang, Siliang, Zhuang, Yueting
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909278605934592
author Gao, Minghe
Li, Juncheng
Fei, Hao
Pang, Liang
Ji, Wei
Wang, Guoming
Lv, Zheqi
Zhang, Wenqiao
Tang, Siliang
Zhuang, Yueting
author_facet Gao, Minghe
Li, Juncheng
Fei, Hao
Pang, Liang
Ji, Wei
Wang, Guoming
Lv, Zheqi
Zhang, Wenqiao
Tang, Siliang
Zhuang, Yueting
contents Visual programming, a modular and generalizable paradigm, integrates different modules and Python operators to solve various vision-language tasks. Unlike end-to-end models that need task-specific data, it advances in performing visual processing and reasoning in an unsupervised manner. Current visual programming methods generate programs in a single pass for each task where the ability to evaluate and optimize based on feedback, unfortunately, is lacking, which consequentially limits their effectiveness for complex, multi-step problems. Drawing inspiration from benders decomposition, we introduce De-fine, a training-free framework that automatically decomposes complex tasks into simpler subtasks and refines programs through auto-feedback. This model-agnostic approach can improve logical reasoning performance by integrating the strengths of multiple models. Our experiments across various visual tasks show that De-fine creates more robust programs. Moreover, viewing each feedback module as an independent agent will yield fresh prospects for the field of agent research.
format Preprint
id arxiv_https___arxiv_org_abs_2311_12890
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle De-fine: Decomposing and Refining Visual Programs with Auto-Feedback
Gao, Minghe
Li, Juncheng
Fei, Hao
Pang, Liang
Ji, Wei
Wang, Guoming
Lv, Zheqi
Zhang, Wenqiao
Tang, Siliang
Zhuang, Yueting
Computer Vision and Pattern Recognition
Visual programming, a modular and generalizable paradigm, integrates different modules and Python operators to solve various vision-language tasks. Unlike end-to-end models that need task-specific data, it advances in performing visual processing and reasoning in an unsupervised manner. Current visual programming methods generate programs in a single pass for each task where the ability to evaluate and optimize based on feedback, unfortunately, is lacking, which consequentially limits their effectiveness for complex, multi-step problems. Drawing inspiration from benders decomposition, we introduce De-fine, a training-free framework that automatically decomposes complex tasks into simpler subtasks and refines programs through auto-feedback. This model-agnostic approach can improve logical reasoning performance by integrating the strengths of multiple models. Our experiments across various visual tasks show that De-fine creates more robust programs. Moreover, viewing each feedback module as an independent agent will yield fresh prospects for the field of agent research.
title De-fine: Decomposing and Refining Visual Programs with Auto-Feedback
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2311.12890