Pre-Trained LLM is a Semantic-Aware and Generalizable Segmentation Booster

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tang, Fenghe, Ma, Wenxin, He, Zhiyang, Tao, Xiaodong, Jiang, Zihang, Zhou, S. Kevin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918067745849344
author Tang, Fenghe
Ma, Wenxin
He, Zhiyang
Tao, Xiaodong
Jiang, Zihang
Zhou, S. Kevin
author_facet Tang, Fenghe
Ma, Wenxin
He, Zhiyang
Tao, Xiaodong
Jiang, Zihang
Zhou, S. Kevin
contents With the advancement of Large Language Model (LLM) for natural language processing, this paper presents an intriguing finding: a frozen pre-trained LLM layer can process visual tokens for medical image segmentation tasks. Specifically, we propose a simple hybrid structure that integrates a pre-trained, frozen LLM layer within the CNN encoder-decoder segmentation framework (LLM4Seg). Surprisingly, this design improves segmentation performance with a minimal increase in trainable parameters across various modalities, including ultrasound, dermoscopy, polypscopy, and CT scans. Our in-depth analysis reveals the potential of transferring LLM's semantic awareness to enhance segmentation tasks, offering both improved global understanding and better local modeling capabilities. The improvement proves robust across different LLMs, validated using LLaMA and DeepSeek.
format Preprint
id arxiv_https___arxiv_org_abs_2506_18034
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Pre-Trained LLM is a Semantic-Aware and Generalizable Segmentation Booster
Tang, Fenghe
Ma, Wenxin
He, Zhiyang
Tao, Xiaodong
Jiang, Zihang
Zhou, S. Kevin
Computer Vision and Pattern Recognition
Artificial Intelligence
Multimedia
With the advancement of Large Language Model (LLM) for natural language processing, this paper presents an intriguing finding: a frozen pre-trained LLM layer can process visual tokens for medical image segmentation tasks. Specifically, we propose a simple hybrid structure that integrates a pre-trained, frozen LLM layer within the CNN encoder-decoder segmentation framework (LLM4Seg). Surprisingly, this design improves segmentation performance with a minimal increase in trainable parameters across various modalities, including ultrasound, dermoscopy, polypscopy, and CT scans. Our in-depth analysis reveals the potential of transferring LLM's semantic awareness to enhance segmentation tasks, offering both improved global understanding and better local modeling capabilities. The improvement proves robust across different LLMs, validated using LLaMA and DeepSeek.
title Pre-Trained LLM is a Semantic-Aware and Generalizable Segmentation Booster
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Multimedia
url https://arxiv.org/abs/2506.18034