Zero-Shot Vision Encoder Grafting via LLM Surrogates

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yue, Kaiyu, Singla, Vasu, Jia, Menglin, Kirchenbauer, John, Qadri, Rifaa, Cai, Zikui, Bhatele, Abhinav, Huang, Furong, Goldstein, Tom
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909717814575104
author Yue, Kaiyu
Singla, Vasu
Jia, Menglin
Kirchenbauer, John
Qadri, Rifaa
Cai, Zikui
Bhatele, Abhinav
Huang, Furong
Goldstein, Tom
author_facet Yue, Kaiyu
Singla, Vasu
Jia, Menglin
Kirchenbauer, John
Qadri, Rifaa
Cai, Zikui
Bhatele, Abhinav
Huang, Furong
Goldstein, Tom
contents Vision language models (VLMs) typically pair a modestly sized vision encoder with a large language model (LLM), e.g., Llama-70B, making the decoder the primary computational burden during training. To reduce costs, a potential promising strategy is to first train the vision encoder using a small language model before transferring it to the large one. We construct small "surrogate models" that share the same embedding space and representation language as the large target LLM by directly inheriting its shallow layers. Vision encoders trained on the surrogate can then be directly transferred to the larger model, a process we call zero-shot grafting -- when plugged directly into the full-size target LLM, the grafted pair surpasses the encoder-surrogate pair and, on some benchmarks, even performs on par with full decoder training with the target LLM. Furthermore, our surrogate training approach reduces overall VLM training costs by ~45% when using Llama-70B as the decoder. The code is at https://github.com/facebookresearch/zero.
format Preprint
id arxiv_https___arxiv_org_abs_2505_22664
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Zero-Shot Vision Encoder Grafting via LLM Surrogates
Yue, Kaiyu
Singla, Vasu
Jia, Menglin
Kirchenbauer, John
Qadri, Rifaa
Cai, Zikui
Bhatele, Abhinav
Huang, Furong
Goldstein, Tom
Computer Vision and Pattern Recognition
Vision language models (VLMs) typically pair a modestly sized vision encoder with a large language model (LLM), e.g., Llama-70B, making the decoder the primary computational burden during training. To reduce costs, a potential promising strategy is to first train the vision encoder using a small language model before transferring it to the large one. We construct small "surrogate models" that share the same embedding space and representation language as the large target LLM by directly inheriting its shallow layers. Vision encoders trained on the surrogate can then be directly transferred to the larger model, a process we call zero-shot grafting -- when plugged directly into the full-size target LLM, the grafted pair surpasses the encoder-surrogate pair and, on some benchmarks, even performs on par with full decoder training with the target LLM. Furthermore, our surrogate training approach reduces overall VLM training costs by ~45% when using Llama-70B as the decoder. The code is at https://github.com/facebookresearch/zero.
title Zero-Shot Vision Encoder Grafting via LLM Surrogates
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.22664