Boosting Salient Object Detection with Knowledge Distillated from Large Foundation Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: He, Miaoyang, Gao, Shuyong, Mok, Tsui Qin, Ge, Weifeng, Zhang, Wengqiang
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913641423437824
author He, Miaoyang
Gao, Shuyong
Mok, Tsui Qin
Ge, Weifeng
Zhang, Wengqiang
author_facet He, Miaoyang
Gao, Shuyong
Mok, Tsui Qin
Ge, Weifeng
Zhang, Wengqiang
contents Salient Object Detection (SOD) aims to identify and segment prominent regions within a scene. Traditional models rely on manually annotated pseudo labels with precise pixel-level accuracy, which is time-consuming. We developed a low-cost, high-precision annotation method by leveraging large foundation models to address the challenges. Specifically, we use a weakly supervised approach to guide large models in generating pseudo-labels through textual prompts. Since large models do not effectively focus on the salient regions of images, we manually annotate a subset of text to fine-tune the model. Based on this approach, which enables precise and rapid generation of pseudo-labels, we introduce a new dataset, BDS-TR. Compared to the previous DUTS-TR dataset, BDS-TR is more prominent in scale and encompasses a wider variety of categories and scenes. This expansion will enhance our model's applicability across a broader range of scenarios and provide a more comprehensive foundational dataset for future SOD research. Additionally, we present an edge decoder based on dynamic upsampling, which focuses on object edges while gradually recovering image feature resolution. Comprehensive experiments on five benchmark datasets demonstrate that our method significantly outperforms state-of-the-art approaches and also surpasses several existing fully-supervised SOD methods. The code and results will be made available.
format Preprint
id arxiv_https___arxiv_org_abs_2501_04582
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Boosting Salient Object Detection with Knowledge Distillated from Large Foundation Models
He, Miaoyang
Gao, Shuyong
Mok, Tsui Qin
Ge, Weifeng
Zhang, Wengqiang
Computer Vision and Pattern Recognition
Salient Object Detection (SOD) aims to identify and segment prominent regions within a scene. Traditional models rely on manually annotated pseudo labels with precise pixel-level accuracy, which is time-consuming. We developed a low-cost, high-precision annotation method by leveraging large foundation models to address the challenges. Specifically, we use a weakly supervised approach to guide large models in generating pseudo-labels through textual prompts. Since large models do not effectively focus on the salient regions of images, we manually annotate a subset of text to fine-tune the model. Based on this approach, which enables precise and rapid generation of pseudo-labels, we introduce a new dataset, BDS-TR. Compared to the previous DUTS-TR dataset, BDS-TR is more prominent in scale and encompasses a wider variety of categories and scenes. This expansion will enhance our model's applicability across a broader range of scenarios and provide a more comprehensive foundational dataset for future SOD research. Additionally, we present an edge decoder based on dynamic upsampling, which focuses on object edges while gradually recovering image feature resolution. Comprehensive experiments on five benchmark datasets demonstrate that our method significantly outperforms state-of-the-art approaches and also surpasses several existing fully-supervised SOD methods. The code and results will be made available.
title Boosting Salient Object Detection with Knowledge Distillated from Large Foundation Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.04582