Vision-Language Controlled Deep Unfolding for Joint Medical Image Restoration and Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Ping, Huang, Zicheng, Wang, Xiangming, Liu, Yungeng, Liang, Bingyu, Zeng, Haijin, Chen, Yongyong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912862933352448
author Chen, Ping
Huang, Zicheng
Wang, Xiangming
Liu, Yungeng
Liang, Bingyu
Zeng, Haijin
Chen, Yongyong
author_facet Chen, Ping
Huang, Zicheng
Wang, Xiangming
Liu, Yungeng
Liang, Bingyu
Zeng, Haijin
Chen, Yongyong
contents We propose VL-DUN, a principled framework for joint All-in-One Medical Image Restoration and Segmentation (AiOMIRS) that bridges the gap between low-level signal recovery and high-level semantic understanding. While standard pipelines treat these tasks in isolation, our core insight is that they are fundamentally synergistic: restoration provides clean anatomical structures to improve segmentation, while semantic priors regularize the restoration process. VL-DUN resolves the sub-optimality of sequential processing through two primary innovations. (1) We formulate AiOMIRS as a unified optimization problem, deriving an interpretable joint unfolding mechanism where restoration and segmentation are mathematically coupled for mutual refinement. (2) We introduce a frequency-aware Mamba mechanism to capture long-range dependencies for global segmentation while preserving the high-frequency textures necessary for restoration. This allows for efficient global context modeling with linear complexity, effectively mitigating the spectral bias of standard architectures. As a pioneering work in the AiOMIRS task, VL-DUN establishes a new state-of-the-art across multi-modal benchmarks, improving PSNR by 0.92 dB and the Dice coefficient by 9.76\%. Our results demonstrate that joint collaborative learning offers a superior, more robust solution for complex clinical workflows compared to isolated task processing. The codes are provided in https://github.com/cipi666/VLDUN.
format Preprint
id arxiv_https___arxiv_org_abs_2601_23103
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Vision-Language Controlled Deep Unfolding for Joint Medical Image Restoration and Segmentation
Chen, Ping
Huang, Zicheng
Wang, Xiangming
Liu, Yungeng
Liang, Bingyu
Zeng, Haijin
Chen, Yongyong
Image and Video Processing
Computer Vision and Pattern Recognition
We propose VL-DUN, a principled framework for joint All-in-One Medical Image Restoration and Segmentation (AiOMIRS) that bridges the gap between low-level signal recovery and high-level semantic understanding. While standard pipelines treat these tasks in isolation, our core insight is that they are fundamentally synergistic: restoration provides clean anatomical structures to improve segmentation, while semantic priors regularize the restoration process. VL-DUN resolves the sub-optimality of sequential processing through two primary innovations. (1) We formulate AiOMIRS as a unified optimization problem, deriving an interpretable joint unfolding mechanism where restoration and segmentation are mathematically coupled for mutual refinement. (2) We introduce a frequency-aware Mamba mechanism to capture long-range dependencies for global segmentation while preserving the high-frequency textures necessary for restoration. This allows for efficient global context modeling with linear complexity, effectively mitigating the spectral bias of standard architectures. As a pioneering work in the AiOMIRS task, VL-DUN establishes a new state-of-the-art across multi-modal benchmarks, improving PSNR by 0.92 dB and the Dice coefficient by 9.76\%. Our results demonstrate that joint collaborative learning offers a superior, more robust solution for complex clinical workflows compared to isolated task processing. The codes are provided in https://github.com/cipi666/VLDUN.
title Vision-Language Controlled Deep Unfolding for Joint Medical Image Restoration and Segmentation
topic Image and Video Processing
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.23103