Staff View: :: Library Catalog

Saved in:

Bibliographic Details
Main Authors:	Wan, Zishuo, Gao, Yu, Pang, Wanyuan, Ding, Dawei
Format:	Preprint
Published:	2025
Subjects:	Computer Vision and Pattern Recognition
Online Access:	https://arxiv.org/abs/2501.03482
Tags:	Add Tag No Tags, Be the first to tag this record!

_version_	1866916553590571008
author	Wan, Zishuo Gao, Yu Pang, Wanyuan Ding, Dawei
author_facet	Wan, Zishuo Gao, Yu Pang, Wanyuan Ding, Dawei
contents	Satisfactory progress has been achieved recently in universal segmentation of CT images. Following the success of vision-language methods, there is a growing trend towards utilizing text prompts and contrastive learning to develop universal segmentation models. However, there exists a significant imbalance in information density between 3D images and text prompts. Moreover, the standard fully connected layer segmentation approach faces significant challenges in handling multiple classes and exhibits poor generalizability. To address these challenges, we propose the VOxel Interacting with LAnguage method (VOILA) for universal CT image segmentation. Initially, we align voxels and language into a shared representation space and classify voxels on the basis of cosine similarity. Subsequently, we develop the Voxel-Language Interaction framework to mitigate the impact of class imbalance caused by foreground-background discrepancies and variations in target volumes. Furthermore, a Complexity-Aware Sampling method is proposed to focus on region hard to segment, achieved by generating pseudo-heatmaps from a trainable Gaussian mixture distribution. Our results indicate the proposed VOILA is capable to achieve improved performance with reduced parameters and computational cost during training. Furthermore, it demonstrates significant generalizability across diverse datasets without additional fine-tuning.
format	Preprint
id	arxiv_https___arxiv_org_abs_2501_03482
institution	arXiv
publishDate	2025
record_format	arxiv
spellingShingle	VOILA: Complexity-Aware Universal Segmentation of CT images by Voxel Interacting with Language Wan, Zishuo Gao, Yu Pang, Wanyuan Ding, Dawei Computer Vision and Pattern Recognition Satisfactory progress has been achieved recently in universal segmentation of CT images. Following the success of vision-language methods, there is a growing trend towards utilizing text prompts and contrastive learning to develop universal segmentation models. However, there exists a significant imbalance in information density between 3D images and text prompts. Moreover, the standard fully connected layer segmentation approach faces significant challenges in handling multiple classes and exhibits poor generalizability. To address these challenges, we propose the VOxel Interacting with LAnguage method (VOILA) for universal CT image segmentation. Initially, we align voxels and language into a shared representation space and classify voxels on the basis of cosine similarity. Subsequently, we develop the Voxel-Language Interaction framework to mitigate the impact of class imbalance caused by foreground-background discrepancies and variations in target volumes. Furthermore, a Complexity-Aware Sampling method is proposed to focus on region hard to segment, achieved by generating pseudo-heatmaps from a trainable Gaussian mixture distribution. Our results indicate the proposed VOILA is capable to achieve improved performance with reduced parameters and computational cost during training. Furthermore, it demonstrates significant generalizability across diverse datasets without additional fine-tuning.
title	VOILA: Complexity-Aware Universal Segmentation of CT images by Voxel Interacting with Language
topic	Computer Vision and Pattern Recognition
url	https://arxiv.org/abs/2501.03482

Similar Items