Visual Program Distillation with Template-Based Augmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shlapentokh-Rothman, Michal, Wang, Yu-Xiong, Hoiem, Derek
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917057422950400
author Shlapentokh-Rothman, Michal
Wang, Yu-Xiong
Hoiem, Derek
author_facet Shlapentokh-Rothman, Michal
Wang, Yu-Xiong
Hoiem, Derek
contents Adapting visual programming or prompting large language models (LLMs) to generate executable code for visual tasks like visual question answering (VQA) for specialized tasks or domains remains challenging due to high annotation and inference costs. We propose a low-cost visual program distillation method that can be used for models with at most 1 billion parameters and requires no human-generated program annotations. We achieve this through synthetic data augmentation based on decoupling programs into higher-level skills, called templates, and their corresponding arguments. Experimental results show that, with a relatively small amount of question/answer data, small language models can generate high-quality specialized visual programs with the added benefit of much faster inference
format Preprint
id arxiv_https___arxiv_org_abs_2412_08564
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Visual Program Distillation with Template-Based Augmentation
Shlapentokh-Rothman, Michal
Wang, Yu-Xiong
Hoiem, Derek
Computer Vision and Pattern Recognition
Computation and Language
Adapting visual programming or prompting large language models (LLMs) to generate executable code for visual tasks like visual question answering (VQA) for specialized tasks or domains remains challenging due to high annotation and inference costs. We propose a low-cost visual program distillation method that can be used for models with at most 1 billion parameters and requires no human-generated program annotations. We achieve this through synthetic data augmentation based on decoupling programs into higher-level skills, called templates, and their corresponding arguments. Experimental results show that, with a relatively small amount of question/answer data, small language models can generate high-quality specialized visual programs with the added benefit of much faster inference
title Visual Program Distillation with Template-Based Augmentation
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2412.08564