Chameleon: A Data-Efficient Generalist for Dense Visual Prediction in the Wild

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Donggyun, Cho, Seongwoong, Kim, Semin, Luo, Chong, Hong, Seunghoon
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917872923574272
author Kim, Donggyun
Cho, Seongwoong
Kim, Semin
Luo, Chong
Hong, Seunghoon
author_facet Kim, Donggyun
Cho, Seongwoong
Kim, Semin
Luo, Chong
Hong, Seunghoon
contents Large language models have evolved data-efficient generalists, benefiting from the universal language interface and large-scale pre-training. However, constructing a data-efficient generalist for dense visual prediction presents a distinct challenge due to the variation in label structures across different tasks. Consequently, generalization to unseen dense prediction tasks in the low-data regime is not straightforward and has received less attention from previous vision generalists. In this study, we explore a universal model that can flexibly adapt to unseen dense label structures with a few examples, enabling it to serve as a data-efficient vision generalist in diverse real-world scenarios. To this end, we base our method on a powerful meta-learning framework and explore several axes to improve its performance and versatility for real-world problems, such as flexible adaptation mechanisms and scalability. We evaluate our model across a spectrum of unseen real-world scenarios where low-shot learning is desirable, including video, 3D, medical, biological, and user-interactive tasks. Equipped with a generic architecture and an effective adaptation mechanism, our model flexibly adapts to all of these tasks with at most 50 labeled images, showcasing a significant advancement over existing data-efficient generalist approaches. Codes are available at https://github.com/GitGyun/chameleon.
format Preprint
id arxiv_https___arxiv_org_abs_2404_18459
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Chameleon: A Data-Efficient Generalist for Dense Visual Prediction in the Wild
Kim, Donggyun
Cho, Seongwoong
Kim, Semin
Luo, Chong
Hong, Seunghoon
Computer Vision and Pattern Recognition
Large language models have evolved data-efficient generalists, benefiting from the universal language interface and large-scale pre-training. However, constructing a data-efficient generalist for dense visual prediction presents a distinct challenge due to the variation in label structures across different tasks. Consequently, generalization to unseen dense prediction tasks in the low-data regime is not straightforward and has received less attention from previous vision generalists. In this study, we explore a universal model that can flexibly adapt to unseen dense label structures with a few examples, enabling it to serve as a data-efficient vision generalist in diverse real-world scenarios. To this end, we base our method on a powerful meta-learning framework and explore several axes to improve its performance and versatility for real-world problems, such as flexible adaptation mechanisms and scalability. We evaluate our model across a spectrum of unseen real-world scenarios where low-shot learning is desirable, including video, 3D, medical, biological, and user-interactive tasks. Equipped with a generic architecture and an effective adaptation mechanism, our model flexibly adapts to all of these tasks with at most 50 labeled images, showcasing a significant advancement over existing data-efficient generalist approaches. Codes are available at https://github.com/GitGyun/chameleon.
title Chameleon: A Data-Efficient Generalist for Dense Visual Prediction in the Wild
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2404.18459