Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chi, Donghwan, Kim, Hyomin, Oh, Yoonjin, Kim, Yongjin, Lee, Donghoon, Jo, Daejin, Kim, Jongmin, Baek, Junyeob, Ahn, Sungjin, Kim, Sungwoong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913145193234432
author Chi, Donghwan
Kim, Hyomin
Oh, Yoonjin
Kim, Yongjin
Lee, Donghoon
Jo, Daejin
Kim, Jongmin
Baek, Junyeob
Ahn, Sungjin
Kim, Sungwoong
author_facet Chi, Donghwan
Kim, Hyomin
Oh, Yoonjin
Kim, Yongjin
Lee, Donghoon
Jo, Daejin
Kim, Jongmin
Baek, Junyeob
Ahn, Sungjin
Kim, Sungwoong
contents Recently, multimodal large language models (MLLMs) have emerged as a key approach in achieving artificial general intelligence. In particular, vision-language MLLMs have been developed to generate not only text but also visual outputs from multimodal inputs. This advancement requires efficient image tokens that LLMs can process effectively both in input and output. However, existing image tokenization methods for MLLMs typically capture only global abstract concepts or uniformly segmented image patches, restricting MLLMs' capability to effectively understand or generate detailed visual content, particularly at the object level. To address this limitation, we propose an object-centric visual tokenizer based on Slot Attention specifically for MLLMs. In particular, based on the Q-Former encoder, diffusion decoder, and residual vector quantization, our proposed discretized slot tokens can encode local visual details while maintaining high-level semantics, and also align with textual data to be integrated seamlessly within a unified next-token prediction framework of LLMs. The resulting Slot-MLLM demonstrates significant performance improvements over baselines with previous visual tokenizers across various vision-language tasks that entail local detailed comprehension and generation. Notably, this work is the first demonstration of the feasibility of object-centric slot attention performed with MLLMs and in-the-wild natural images.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17726
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM
Chi, Donghwan
Kim, Hyomin
Oh, Yoonjin
Kim, Yongjin
Lee, Donghoon
Jo, Daejin
Kim, Jongmin
Baek, Junyeob
Ahn, Sungjin
Kim, Sungwoong
Computer Vision and Pattern Recognition
Artificial Intelligence
Recently, multimodal large language models (MLLMs) have emerged as a key approach in achieving artificial general intelligence. In particular, vision-language MLLMs have been developed to generate not only text but also visual outputs from multimodal inputs. This advancement requires efficient image tokens that LLMs can process effectively both in input and output. However, existing image tokenization methods for MLLMs typically capture only global abstract concepts or uniformly segmented image patches, restricting MLLMs' capability to effectively understand or generate detailed visual content, particularly at the object level. To address this limitation, we propose an object-centric visual tokenizer based on Slot Attention specifically for MLLMs. In particular, based on the Q-Former encoder, diffusion decoder, and residual vector quantization, our proposed discretized slot tokens can encode local visual details while maintaining high-level semantics, and also align with textual data to be integrated seamlessly within a unified next-token prediction framework of LLMs. The resulting Slot-MLLM demonstrates significant performance improvements over baselines with previous visual tokenizers across various vision-language tasks that entail local detailed comprehension and generation. Notably, this work is the first demonstration of the feasibility of object-centric slot attention performed with MLLMs and in-the-wild natural images.
title Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2505.17726