SegDINO: An Efficient Design for Medical and Natural Image Segmentation with DINO-V3

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Sicheng, Wang, Hongqiu, Xing, Zhaohu, Chen, Sixiang, Zhu, Lei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916927850414080
author Yang, Sicheng
Wang, Hongqiu
Xing, Zhaohu
Chen, Sixiang
Zhu, Lei
author_facet Yang, Sicheng
Wang, Hongqiu
Xing, Zhaohu
Chen, Sixiang
Zhu, Lei
contents The DINO family of self-supervised vision models has shown remarkable transferability, yet effectively adapting their representations for segmentation remains challenging. Existing approaches often rely on heavy decoders with multi-scale fusion or complex upsampling, which introduce substantial parameter overhead and computational cost. In this work, we propose SegDINO, an efficient segmentation framework that couples a frozen DINOv3 backbone with a lightweight decoder. SegDINO extracts multi-level features from the pretrained encoder, aligns them to a common resolution and channel width, and utilizes a lightweight MLP head to directly predict segmentation masks. This design minimizes trainable parameters while preserving the representational power of foundation features. Extensive experiments across six benchmarks, including three medical datasets (TN3K, Kvasir-SEG, ISIC) and three natural image datasets (MSD, VMD-D, ViSha), demonstrate that SegDINO consistently achieves state-of-the-art performance compared to existing methods. Code is available at https://github.com/script-Yang/SegDINO.
format Preprint
id arxiv_https___arxiv_org_abs_2509_00833
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SegDINO: An Efficient Design for Medical and Natural Image Segmentation with DINO-V3
Yang, Sicheng
Wang, Hongqiu
Xing, Zhaohu
Chen, Sixiang
Zhu, Lei
Computer Vision and Pattern Recognition
The DINO family of self-supervised vision models has shown remarkable transferability, yet effectively adapting their representations for segmentation remains challenging. Existing approaches often rely on heavy decoders with multi-scale fusion or complex upsampling, which introduce substantial parameter overhead and computational cost. In this work, we propose SegDINO, an efficient segmentation framework that couples a frozen DINOv3 backbone with a lightweight decoder. SegDINO extracts multi-level features from the pretrained encoder, aligns them to a common resolution and channel width, and utilizes a lightweight MLP head to directly predict segmentation masks. This design minimizes trainable parameters while preserving the representational power of foundation features. Extensive experiments across six benchmarks, including three medical datasets (TN3K, Kvasir-SEG, ISIC) and three natural image datasets (MSD, VMD-D, ViSha), demonstrate that SegDINO consistently achieves state-of-the-art performance compared to existing methods. Code is available at https://github.com/script-Yang/SegDINO.
title SegDINO: An Efficient Design for Medical and Natural Image Segmentation with DINO-V3
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.00833