Learning to Edit Visual Programs with Self-Supervision

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jones, R. Kenny, Zhang, Renhao, Ganeshan, Aditya, Ritchie, Daniel
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909375060246528
author Jones, R. Kenny
Zhang, Renhao
Ganeshan, Aditya
Ritchie, Daniel
author_facet Jones, R. Kenny
Zhang, Renhao
Ganeshan, Aditya
Ritchie, Daniel
contents We design a system that learns how to edit visual programs. Our edit network consumes a complete input program and a visual target. From this input, we task our network with predicting a local edit operation that could be applied to the input program to improve its similarity to the target. In order to apply this scheme for domains that lack program annotations, we develop a self-supervised learning approach that integrates this edit network into a bootstrapped finetuning loop along with a network that predicts entire programs in one-shot. Our joint finetuning scheme, when coupled with an inference procedure that initializes a population from the one-shot model and evolves members of this population with the edit network, helps to infer more accurate visual programs. Over multiple domains, we experimentally compare our method against the alternative of using only the one-shot model, and find that even under equal search-time budgets, our editing-based paradigm provides significant advantages.
format Preprint
id arxiv_https___arxiv_org_abs_2406_02383
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Learning to Edit Visual Programs with Self-Supervision
Jones, R. Kenny
Zhang, Renhao
Ganeshan, Aditya
Ritchie, Daniel
Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
Machine Learning
We design a system that learns how to edit visual programs. Our edit network consumes a complete input program and a visual target. From this input, we task our network with predicting a local edit operation that could be applied to the input program to improve its similarity to the target. In order to apply this scheme for domains that lack program annotations, we develop a self-supervised learning approach that integrates this edit network into a bootstrapped finetuning loop along with a network that predicts entire programs in one-shot. Our joint finetuning scheme, when coupled with an inference procedure that initializes a population from the one-shot model and evolves members of this population with the edit network, helps to infer more accurate visual programs. Over multiple domains, we experimentally compare our method against the alternative of using only the one-shot model, and find that even under equal search-time budgets, our editing-based paradigm provides significant advantages.
title Learning to Edit Visual Programs with Self-Supervision
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
Machine Learning
url https://arxiv.org/abs/2406.02383