JVLGS: Joint Vision-Language Gas Leak Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Xinlong, Pang, Qixiang, Du, Shan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912556082266112
author Zhao, Xinlong
Pang, Qixiang
Du, Shan
author_facet Zhao, Xinlong
Pang, Qixiang
Du, Shan
contents Gas leaks pose serious threats to human health and contribute significantly to atmospheric pollution, drawing increasing public concern. However, the lack of effective detection methods hampers timely and accurate identification of gas leaks. While some vision-based techniques leverage infrared videos for leak detection, the blurry and non-rigid nature of gas clouds often limits their effectiveness. To address these challenges, we propose a novel framework called Joint Vision-Language Gas leak Segmentation (JVLGS), which integrates the complementary strengths of visual and textual modalities to enhance gas leak representation and segmentation. Recognizing that gas leaks are sporadic and many video frames may contain no leak at all, our method incorporates a post-processing step to reduce false positives caused by noise and non-target objects, an issue that affects many existing approaches. Extensive experiments conducted across diverse scenarios show that JVLGS significantly outperforms state-of-the-art gas leak segmentation methods. We evaluate our model under both supervised and few-shot learning settings, and it consistently achieves strong performance in both, whereas competing methods tend to perform well in only one setting or poorly in both. Code available at: https://github.com/GeekEagle/JVLGS
format Preprint
id arxiv_https___arxiv_org_abs_2508_19485
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle JVLGS: Joint Vision-Language Gas Leak Segmentation
Zhao, Xinlong
Pang, Qixiang
Du, Shan
Computer Vision and Pattern Recognition
68T45 (Primary), 68T07 (Secondary)
I.2.10; I.4.6
Gas leaks pose serious threats to human health and contribute significantly to atmospheric pollution, drawing increasing public concern. However, the lack of effective detection methods hampers timely and accurate identification of gas leaks. While some vision-based techniques leverage infrared videos for leak detection, the blurry and non-rigid nature of gas clouds often limits their effectiveness. To address these challenges, we propose a novel framework called Joint Vision-Language Gas leak Segmentation (JVLGS), which integrates the complementary strengths of visual and textual modalities to enhance gas leak representation and segmentation. Recognizing that gas leaks are sporadic and many video frames may contain no leak at all, our method incorporates a post-processing step to reduce false positives caused by noise and non-target objects, an issue that affects many existing approaches. Extensive experiments conducted across diverse scenarios show that JVLGS significantly outperforms state-of-the-art gas leak segmentation methods. We evaluate our model under both supervised and few-shot learning settings, and it consistently achieves strong performance in both, whereas competing methods tend to perform well in only one setting or poorly in both. Code available at: https://github.com/GeekEagle/JVLGS
title JVLGS: Joint Vision-Language Gas Leak Segmentation
topic Computer Vision and Pattern Recognition
68T45 (Primary), 68T07 (Secondary)
I.2.10; I.4.6
url https://arxiv.org/abs/2508.19485