Self-Supervised Multi-Scale Transformer with Attention-Guided Fusion for Efficient Crack Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kyem, Blessing Agyei, Asamoah, Joshua Kofi, Denteh, Eugene, Danyo, Andrews, Aboah, Armstrong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908606633345024
author Kyem, Blessing Agyei
Asamoah, Joshua Kofi
Denteh, Eugene
Danyo, Andrews
Aboah, Armstrong
author_facet Kyem, Blessing Agyei
Asamoah, Joshua Kofi
Denteh, Eugene
Danyo, Andrews
Aboah, Armstrong
contents Pavement crack detection has long depended on costly and time-intensive pixel-level annotations, which limit its scalability for large-scale infrastructure monitoring. To overcome this barrier, this paper examines the feasibility of achieving effective pixel-level crack segmentation entirely without manual annotations. Building on this objective, a fully self-supervised framework, Crack-Segmenter, is developed, integrating three complementary modules: the Scale-Adaptive Embedder (SAE) for robust multi-scale feature extraction, the Directional Attention Transformer (DAT) for maintaining linear crack continuity, and the Attention-Guided Fusion (AGF) module for adaptive feature integration. Through evaluations on ten public datasets, Crack-Segmenter consistently outperforms 13 state-of-the-art supervised methods across all major metrics, including mean Intersection over Union (mIoU), Dice score, XOR, and Hausdorff Distance (HD). These findings demonstrate that annotation-free crack detection is not only feasible but also superior, enabling transportation agencies and infrastructure managers to conduct scalable and cost-effective monitoring. This work advances self-supervised learning and motivates pavement cracks detection research.
format Preprint
id arxiv_https___arxiv_org_abs_2510_10378
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Self-Supervised Multi-Scale Transformer with Attention-Guided Fusion for Efficient Crack Detection
Kyem, Blessing Agyei
Asamoah, Joshua Kofi
Denteh, Eugene
Danyo, Andrews
Aboah, Armstrong
Computer Vision and Pattern Recognition
Pavement crack detection has long depended on costly and time-intensive pixel-level annotations, which limit its scalability for large-scale infrastructure monitoring. To overcome this barrier, this paper examines the feasibility of achieving effective pixel-level crack segmentation entirely without manual annotations. Building on this objective, a fully self-supervised framework, Crack-Segmenter, is developed, integrating three complementary modules: the Scale-Adaptive Embedder (SAE) for robust multi-scale feature extraction, the Directional Attention Transformer (DAT) for maintaining linear crack continuity, and the Attention-Guided Fusion (AGF) module for adaptive feature integration. Through evaluations on ten public datasets, Crack-Segmenter consistently outperforms 13 state-of-the-art supervised methods across all major metrics, including mean Intersection over Union (mIoU), Dice score, XOR, and Hausdorff Distance (HD). These findings demonstrate that annotation-free crack detection is not only feasible but also superior, enabling transportation agencies and infrastructure managers to conduct scalable and cost-effective monitoring. This work advances self-supervised learning and motivates pavement cracks detection research.
title Self-Supervised Multi-Scale Transformer with Attention-Guided Fusion for Efficient Crack Detection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.10378