Towards Semantic Equivalence of Tokenization in Multimodal LLM

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Shengqiong, Fei, Hao, Li, Xiangtai, Ji, Jiayi, Zhang, Hanwang, Chua, Tat-Seng, Yan, Shuicheng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909511090962432
author Wu, Shengqiong
Fei, Hao
Li, Xiangtai
Ji, Jiayi
Zhang, Hanwang
Chua, Tat-Seng
Yan, Shuicheng
author_facet Wu, Shengqiong
Fei, Hao
Li, Xiangtai
Ji, Jiayi
Zhang, Hanwang
Chua, Tat-Seng
Yan, Shuicheng
contents Multimodal Large Language Models (MLLMs) have demonstrated exceptional capabilities in processing vision-language tasks. One of the crux of MLLMs lies in vision tokenization, which involves efficiently transforming input visual signals into feature representations that are most beneficial for LLMs. However, existing vision tokenizers, essential for semantic alignment between vision and language, remain problematic. Existing methods aggressively fragment visual input, corrupting the visual semantic integrity. To address this, this paper proposes a novel dynamic Semantic-Equivalent Vision Tokenizer (SeTok), which groups visual features into semantic units via a dynamic clustering algorithm, flexibly determining the number of tokens based on image complexity. The resulting vision tokens effectively preserve semantic integrity and capture both low-frequency and high-frequency visual features. The proposed MLLM (Setokim) equipped with SeTok significantly demonstrates superior performance across various tasks, as evidenced by our experimental results. The project page is at https://chocowu.github.io/SeTok-web/.
format Preprint
id arxiv_https___arxiv_org_abs_2406_05127
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Towards Semantic Equivalence of Tokenization in Multimodal LLM
Wu, Shengqiong
Fei, Hao
Li, Xiangtai
Ji, Jiayi
Zhang, Hanwang
Chua, Tat-Seng
Yan, Shuicheng
Computer Vision and Pattern Recognition
Multimodal Large Language Models (MLLMs) have demonstrated exceptional capabilities in processing vision-language tasks. One of the crux of MLLMs lies in vision tokenization, which involves efficiently transforming input visual signals into feature representations that are most beneficial for LLMs. However, existing vision tokenizers, essential for semantic alignment between vision and language, remain problematic. Existing methods aggressively fragment visual input, corrupting the visual semantic integrity. To address this, this paper proposes a novel dynamic Semantic-Equivalent Vision Tokenizer (SeTok), which groups visual features into semantic units via a dynamic clustering algorithm, flexibly determining the number of tokens based on image complexity. The resulting vision tokens effectively preserve semantic integrity and capture both low-frequency and high-frequency visual features. The proposed MLLM (Setokim) equipped with SeTok significantly demonstrates superior performance across various tasks, as evidenced by our experimental results. The project page is at https://chocowu.github.io/SeTok-web/.
title Towards Semantic Equivalence of Tokenization in Multimodal LLM
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2406.05127