Unified Medical Image Tokenizer for Autoregressive Synthesis and Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ma, Chenglong, Ji, Yuanfeng, Ye, Jin, Li, Zilong, Wang, Chenhui, Ning, Junzhi, Li, Wei, Liu, Lihao, Guo, Qiushan, Li, Tianbin, He, Junjun, Shan, Hongming
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914436321640448
author Ma, Chenglong
Ji, Yuanfeng
Ye, Jin
Li, Zilong
Wang, Chenhui
Ning, Junzhi
Li, Wei
Liu, Lihao
Guo, Qiushan
Li, Tianbin
He, Junjun
Shan, Hongming
author_facet Ma, Chenglong
Ji, Yuanfeng
Ye, Jin
Li, Zilong
Wang, Chenhui
Ning, Junzhi
Li, Wei
Liu, Lihao
Guo, Qiushan
Li, Tianbin
He, Junjun
Shan, Hongming
contents Autoregressive modeling has driven major advances in multimodal AI, yet its application to medical imaging remains constrained by the absence of a unified image tokenizer that simultaneously preserves fine-grained anatomical structures and rich clinical semantics across heterogeneous modalities. Existing approaches jointly optimize image reconstruction and textual semantic objectives, relying on large-scale image-caption pairs and are prone to gradient interference. This is ill-suited for the medical domain where paired data are scarce and abundant unpaired images remain unexploited. This work identifies these issues in building unified medical image tokenizers, and introduces a principled two-stage training framework using visual representation as a bridge to address them. The propose visual representation alignment stage enables the utilization of large-scale unpaired medical images to ensure reconstruction fidelity and establish foundational semantics, alleviating the interference and better preparing for the second stage where fine-grained textual semantics are injected using image-text pairs. The resulting tokenizer, MedITok, is trained on over 33 million medical images spanning 9 modalities and 2 million image-text pairs. MedITok achieves state-of-the-art performance on 30+ benchmarks spanning 9 imaging modalities and 4 task families. It further enables autoregressive modeling for diagnostic and generative applications, serving as a scalable component for future multimodal models with unified synthesis and understanding capabilities in the medical domain. Project page: https://github.com/Masaaki-75/meditok
format Preprint
id arxiv_https___arxiv_org_abs_2505_19225
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Unified Medical Image Tokenizer for Autoregressive Synthesis and Understanding
Ma, Chenglong
Ji, Yuanfeng
Ye, Jin
Li, Zilong
Wang, Chenhui
Ning, Junzhi
Li, Wei
Liu, Lihao
Guo, Qiushan
Li, Tianbin
He, Junjun
Shan, Hongming
Image and Video Processing
Computer Vision and Pattern Recognition
Autoregressive modeling has driven major advances in multimodal AI, yet its application to medical imaging remains constrained by the absence of a unified image tokenizer that simultaneously preserves fine-grained anatomical structures and rich clinical semantics across heterogeneous modalities. Existing approaches jointly optimize image reconstruction and textual semantic objectives, relying on large-scale image-caption pairs and are prone to gradient interference. This is ill-suited for the medical domain where paired data are scarce and abundant unpaired images remain unexploited. This work identifies these issues in building unified medical image tokenizers, and introduces a principled two-stage training framework using visual representation as a bridge to address them. The propose visual representation alignment stage enables the utilization of large-scale unpaired medical images to ensure reconstruction fidelity and establish foundational semantics, alleviating the interference and better preparing for the second stage where fine-grained textual semantics are injected using image-text pairs. The resulting tokenizer, MedITok, is trained on over 33 million medical images spanning 9 modalities and 2 million image-text pairs. MedITok achieves state-of-the-art performance on 30+ benchmarks spanning 9 imaging modalities and 4 task families. It further enables autoregressive modeling for diagnostic and generative applications, serving as a scalable component for future multimodal models with unified synthesis and understanding capabilities in the medical domain. Project page: https://github.com/Masaaki-75/meditok
title Unified Medical Image Tokenizer for Autoregressive Synthesis and Understanding
topic Image and Video Processing
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.19225