Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Su, Yongyi, Zhang, Haojie, Li, Shijie, Liu, Nanqing, Liao, Jingyi, Pan, Junyi, Liu, Yuan, Xing, Xiaofen, Sun, Chong, Li, Chen, Chen, Nancy F., Yan, Shuicheng, Yang, Xulei, Xu, Xun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914071691919360
author Su, Yongyi
Zhang, Haojie
Li, Shijie
Liu, Nanqing
Liao, Jingyi
Pan, Junyi
Liu, Yuan
Xing, Xiaofen
Sun, Chong
Li, Chen
Chen, Nancy F.
Yan, Shuicheng
Yang, Xulei
Xu, Xun
author_facet Su, Yongyi
Zhang, Haojie
Li, Shijie
Liu, Nanqing
Liao, Jingyi
Pan, Junyi
Liu, Yuan
Xing, Xiaofen
Sun, Chong
Li, Chen
Chen, Nancy F.
Yan, Shuicheng
Yang, Xulei
Xu, Xun
contents Multimodal large language models (MLLMs) have advanced rapidly in recent years. However, existing approaches for vision tasks often rely on indirect representations, such as generating coordinates as text for detection, which limits performance and prevents dense prediction tasks like segmentation. To overcome these challenges, we introduce Patch-as-Decodable Token (PaDT), a unified paradigm that enables MLLMs to directly generate both textual and diverse visual outputs. Central to PaDT are Visual Reference Tokens (VRTs), derived from visual patch embeddings of query images and interleaved seamlessly with LLM's output textual tokens. A lightweight decoder then transforms LLM's outputs into detection, segmentation, and grounding predictions. Unlike prior methods, PaDT processes VRTs independently at each forward pass and dynamically expands the embedding table, thus improving localization and differentiation among similar objects. We further tailor a training strategy for PaDT by randomly selecting VRTs for supervised fine-tuning and introducing a robust per-token cross-entropy loss. Our empirical studies across four visual perception and understanding tasks suggest PaDT consistently achieving state-of-the-art performance, even compared with significantly larger MLLM models. The code is available at https://github.com/Gorilla-Lab-SCUT/PaDT.
format Preprint
id arxiv_https___arxiv_org_abs_2510_01954
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs
Su, Yongyi
Zhang, Haojie
Li, Shijie
Liu, Nanqing
Liao, Jingyi
Pan, Junyi
Liu, Yuan
Xing, Xiaofen
Sun, Chong
Li, Chen
Chen, Nancy F.
Yan, Shuicheng
Yang, Xulei
Xu, Xun
Computer Vision and Pattern Recognition
Multimodal large language models (MLLMs) have advanced rapidly in recent years. However, existing approaches for vision tasks often rely on indirect representations, such as generating coordinates as text for detection, which limits performance and prevents dense prediction tasks like segmentation. To overcome these challenges, we introduce Patch-as-Decodable Token (PaDT), a unified paradigm that enables MLLMs to directly generate both textual and diverse visual outputs. Central to PaDT are Visual Reference Tokens (VRTs), derived from visual patch embeddings of query images and interleaved seamlessly with LLM's output textual tokens. A lightweight decoder then transforms LLM's outputs into detection, segmentation, and grounding predictions. Unlike prior methods, PaDT processes VRTs independently at each forward pass and dynamically expands the embedding table, thus improving localization and differentiation among similar objects. We further tailor a training strategy for PaDT by randomly selecting VRTs for supervised fine-tuning and introducing a robust per-token cross-entropy loss. Our empirical studies across four visual perception and understanding tasks suggest PaDT consistently achieving state-of-the-art performance, even compared with significantly larger MLLM models. The code is available at https://github.com/Gorilla-Lab-SCUT/PaDT.
title Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.01954