Scaling Capability in Token Space: An Analysis of Large Vision Language Model

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Li, Tenghui, Zhou, Guoxu, Zhao, Xuyang, Zhao, Qibin
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911341750517760
author Li, Tenghui
Zhou, Guoxu
Zhao, Xuyang
Zhao, Qibin
author_facet Li, Tenghui
Zhou, Guoxu
Zhao, Xuyang
Zhao, Qibin
contents Large language models have demonstrated predictable scaling behaviors with respect to model parameters and training data. This study investigates whether a similar scaling relationship exist for vision-language models with respect to the number of vision tokens. A mathematical framework is developed to characterize a relationship between vision token number and the expected divergence of distance between vision-referencing sequences. The theoretical analysis reveals two distinct scaling regimes: sublinear scaling for less vision tokens and linear scaling for more vision tokens. This aligns with model performance relationships of the form \(S(n) \approx c / n^{α(n)}\), where the scaling exponent relates to the correlation structure between vision token representations. Empirical validations across multiple vision-language benchmarks show that model performance matches the prediction from scaling relationship. The findings contribute to understanding vision token scaling in transformers through a theoretical framework that complements empirical observations.
format Preprint
id arxiv_https___arxiv_org_abs_2412_18387
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Scaling Capability in Token Space: An Analysis of Large Vision Language Model
Li, Tenghui
Zhou, Guoxu
Zhao, Xuyang
Zhao, Qibin
Artificial Intelligence
Machine Learning
Large language models have demonstrated predictable scaling behaviors with respect to model parameters and training data. This study investigates whether a similar scaling relationship exist for vision-language models with respect to the number of vision tokens. A mathematical framework is developed to characterize a relationship between vision token number and the expected divergence of distance between vision-referencing sequences. The theoretical analysis reveals two distinct scaling regimes: sublinear scaling for less vision tokens and linear scaling for more vision tokens. This aligns with model performance relationships of the form \(S(n) \approx c / n^{α(n)}\), where the scaling exponent relates to the correlation structure between vision token representations. Empirical validations across multiple vision-language benchmarks show that model performance matches the prediction from scaling relationship. The findings contribute to understanding vision token scaling in transformers through a theoretical framework that complements empirical observations.
title Scaling Capability in Token Space: An Analysis of Large Vision Language Model
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2412.18387