Saved in:
Bibliographic Details
Main Authors: Torabi, Ali, Gaihre, Sanjog, Majeed, Yaqoob
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2511.19765
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914170000113664
author Torabi, Ali
Gaihre, Sanjog
Majeed, Yaqoob
author_facet Torabi, Ali
Gaihre, Sanjog
Majeed, Yaqoob
contents Weakly supervised semantic segmentation (WSSS) must learn dense masks from noisy, under-specified cues. We revisit the SegFormer decoder and show that three small, synergistic changes make weak supervision markedly more effective-without altering the MiT backbone or relying on heavy post-processing. Our method, CrispFormer, augments the decoder with: (1) a boundary branch that supervises thin object contours using a lightweight edge head and a boundary-aware loss; (2) an uncertainty-guided refiner that predicts per-pixel aleatoric uncertainty and uses it to weight losses and gate a residual correction of the segmentation logits; and (3) a dynamic multi-scale fusion layer that replaces static concatenation with spatial softmax gating over multi-resolution features, optionally modulated by uncertainty. The result is a single-pass model that preserves crisp boundaries, selects appropriate scales per location, and resists label noise from weak cues. Integrated into a standard WSSS pipeline (seed, student, and EMA relabeling), CrispFormer consistently improves boundary F-score, small-object recall, and mIoU over SegFormer baselines trained on the same seeds, while adding minimal compute. Our decoder-centric formulation is simple to implement, broadly compatible with existing SegFormer variants, and offers a reproducible path to higher-fidelity masks from image-level supervision.
format Preprint
id arxiv_https___arxiv_org_abs_2511_19765
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Lightweight Transformer Framework for Weakly Supervised Semantic Segmentation
Torabi, Ali
Gaihre, Sanjog
Majeed, Yaqoob
Computer Vision and Pattern Recognition
Weakly supervised semantic segmentation (WSSS) must learn dense masks from noisy, under-specified cues. We revisit the SegFormer decoder and show that three small, synergistic changes make weak supervision markedly more effective-without altering the MiT backbone or relying on heavy post-processing. Our method, CrispFormer, augments the decoder with: (1) a boundary branch that supervises thin object contours using a lightweight edge head and a boundary-aware loss; (2) an uncertainty-guided refiner that predicts per-pixel aleatoric uncertainty and uses it to weight losses and gate a residual correction of the segmentation logits; and (3) a dynamic multi-scale fusion layer that replaces static concatenation with spatial softmax gating over multi-resolution features, optionally modulated by uncertainty. The result is a single-pass model that preserves crisp boundaries, selects appropriate scales per location, and resists label noise from weak cues. Integrated into a standard WSSS pipeline (seed, student, and EMA relabeling), CrispFormer consistently improves boundary F-score, small-object recall, and mIoU over SegFormer baselines trained on the same seeds, while adding minimal compute. Our decoder-centric formulation is simple to implement, broadly compatible with existing SegFormer variants, and offers a reproducible path to higher-fidelity masks from image-level supervision.
title Lightweight Transformer Framework for Weakly Supervised Semantic Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.19765