The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lei, Weixian, Wang, Jiacong, Wang, Haochen, Li, Xiangtai, Liew, Jun Hao, Feng, Jiashi, Huang, Zilong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916689146281984
author Lei, Weixian
Wang, Jiacong
Wang, Haochen
Li, Xiangtai
Liew, Jun Hao
Feng, Jiashi
Huang, Zilong
author_facet Lei, Weixian
Wang, Jiacong
Wang, Haochen
Li, Xiangtai
Liew, Jun Hao
Feng, Jiashi
Huang, Zilong
contents This paper introduces SAIL, a single transformer unified multimodal large language model (MLLM) that integrates raw pixel encoding and language decoding within a singular architecture. Unlike existing modular MLLMs, which rely on a pre-trained vision transformer (ViT), SAIL eliminates the need for a separate vision encoder, presenting a more minimalist architecture design. Instead of introducing novel architectural components, SAIL adapts mix-attention mechanisms and multimodal positional encodings to better align with the distinct characteristics of visual and textual modalities. We systematically compare SAIL's properties-including scalability, cross-modal information flow patterns, and visual representation capabilities-with those of modular MLLMs. By scaling both training data and model size, SAIL achieves performance comparable to modular MLLMs. Notably, the removal of pretrained ViT components enhances SAIL's scalability and results in significantly different cross-modal information flow patterns. Moreover, SAIL demonstrates strong visual representation capabilities, achieving results on par with ViT-22B in vision tasks such as semantic segmentation. Code and models are available at https://github.com/bytedance/SAIL.
format Preprint
id arxiv_https___arxiv_org_abs_2504_10462
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer
Lei, Weixian
Wang, Jiacong
Wang, Haochen
Li, Xiangtai
Liew, Jun Hao
Feng, Jiashi
Huang, Zilong
Computer Vision and Pattern Recognition
This paper introduces SAIL, a single transformer unified multimodal large language model (MLLM) that integrates raw pixel encoding and language decoding within a singular architecture. Unlike existing modular MLLMs, which rely on a pre-trained vision transformer (ViT), SAIL eliminates the need for a separate vision encoder, presenting a more minimalist architecture design. Instead of introducing novel architectural components, SAIL adapts mix-attention mechanisms and multimodal positional encodings to better align with the distinct characteristics of visual and textual modalities. We systematically compare SAIL's properties-including scalability, cross-modal information flow patterns, and visual representation capabilities-with those of modular MLLMs. By scaling both training data and model size, SAIL achieves performance comparable to modular MLLMs. Notably, the removal of pretrained ViT components enhances SAIL's scalability and results in significantly different cross-modal information flow patterns. Moreover, SAIL demonstrates strong visual representation capabilities, achieving results on par with ViT-22B in vision tasks such as semantic segmentation. Code and models are available at https://github.com/bytedance/SAIL.
title The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.10462