CLIP-Optimized Multimodal Image Enhancement via ISP-CNN Fusion for Coal Mine IoVT under Uneven Illumination

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Shuai, Zhang, Shihao, Wu, Jiaqi, Tian, Zijian, Chen, Wei, Jin, Tongzhu, Xue, Miaomiao, Wang, Zehua, Yu, Fei Richard, Leung, Victor C. M.
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916632849285120
author Wang, Shuai
Zhang, Shihao
Wu, Jiaqi
Tian, Zijian
Chen, Wei
Jin, Tongzhu
Xue, Miaomiao
Wang, Zehua
Yu, Fei Richard
Leung, Victor C. M.
author_facet Wang, Shuai
Zhang, Shihao
Wu, Jiaqi
Tian, Zijian
Chen, Wei
Jin, Tongzhu
Xue, Miaomiao
Wang, Zehua
Yu, Fei Richard
Leung, Victor C. M.
contents Clear monitoring images are crucial for the safe operation of coal mine Internet of Video Things (IoVT) systems. However, low illumination and uneven brightness in underground environments significantly degrade image quality, posing challenges for enhancement methods that often rely on difficult-to-obtain paired reference images. Additionally, there is a trade-off between enhancement performance and computational efficiency on edge devices within IoVT systems.To address these issues, we propose a multimodal image enhancement method tailored for coal mine IoVT, utilizing an ISP-CNN fusion architecture optimized for uneven illumination. This two-stage strategy combines global enhancement with detail optimization, effectively improving image quality, especially in poorly lit areas. A CLIP-based multimodal iterative optimization allows for unsupervised training of the enhancement algorithm. By integrating traditional image signal processing (ISP) with convolutional neural networks (CNN), our approach reduces computational complexity while maintaining high performance, making it suitable for real-time deployment on edge devices.Experimental results demonstrate that our method effectively mitigates uneven brightness and enhances key image quality metrics, with PSNR improvements of 2.9%-4.9%, SSIM by 4.3%-11.4%, and VIF by 4.9%-17.8% compared to seven state-of-the-art algorithms. Simulated coal mine monitoring scenarios validate our method's ability to balance performance and computational demands, facilitating real-time enhancement and supporting safer mining operations.
format Preprint
id arxiv_https___arxiv_org_abs_2502_19450
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CLIP-Optimized Multimodal Image Enhancement via ISP-CNN Fusion for Coal Mine IoVT under Uneven Illumination
Wang, Shuai
Zhang, Shihao
Wu, Jiaqi
Tian, Zijian
Chen, Wei
Jin, Tongzhu
Xue, Miaomiao
Wang, Zehua
Yu, Fei Richard
Leung, Victor C. M.
Computer Vision and Pattern Recognition
Clear monitoring images are crucial for the safe operation of coal mine Internet of Video Things (IoVT) systems. However, low illumination and uneven brightness in underground environments significantly degrade image quality, posing challenges for enhancement methods that often rely on difficult-to-obtain paired reference images. Additionally, there is a trade-off between enhancement performance and computational efficiency on edge devices within IoVT systems.To address these issues, we propose a multimodal image enhancement method tailored for coal mine IoVT, utilizing an ISP-CNN fusion architecture optimized for uneven illumination. This two-stage strategy combines global enhancement with detail optimization, effectively improving image quality, especially in poorly lit areas. A CLIP-based multimodal iterative optimization allows for unsupervised training of the enhancement algorithm. By integrating traditional image signal processing (ISP) with convolutional neural networks (CNN), our approach reduces computational complexity while maintaining high performance, making it suitable for real-time deployment on edge devices.Experimental results demonstrate that our method effectively mitigates uneven brightness and enhances key image quality metrics, with PSNR improvements of 2.9%-4.9%, SSIM by 4.3%-11.4%, and VIF by 4.9%-17.8% compared to seven state-of-the-art algorithms. Simulated coal mine monitoring scenarios validate our method's ability to balance performance and computational demands, facilitating real-time enhancement and supporting safer mining operations.
title CLIP-Optimized Multimodal Image Enhancement via ISP-CNN Fusion for Coal Mine IoVT under Uneven Illumination
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2502.19450