M2Retinexformer: Multi-Modal Retinexformer for Low-Light Image Enhancement

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Aboelwafa, Youssef, Elmongui, Hicham G., Torki, Marwan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911677041082368
author Aboelwafa, Youssef
Elmongui, Hicham G.
Torki, Marwan
author_facet Aboelwafa, Youssef
Elmongui, Hicham G.
Torki, Marwan
contents Low-light image enhancement is challenging due to complex degradations, including amplified noise, artifacts, and color distortion. While Retinex-based deep learning methods have achieved promising results, they primarily rely on single-modality RGB information. We propose M2Retinexformer (Multi-Modal Retinexformer), a novel framework that extends Retinexformer by incorporating depth cues, luminance priors, and semantic features within a progressive refinement pipeline. Depth provides geometric context that is invariant to lighting variations, while luminance and semantic features offer explicit guidance on brightness distribution and scene understanding. Modalities are extracted at multiple scales and fused through cross-attention, with adaptive gating dynamically balancing illumination-guided self-attention and cross-attention based on the reliability of auxiliary cues. Evaluations on the LOL, SID, SMID, and SDSD benchmarks demonstrate overall improvements over Retinexformer and recent state-of-the-art methods. Code and pretrained weights are available at https://github.com/YoussefAboelwafa/M2Retinexformer
format Preprint
id arxiv_https___arxiv_org_abs_2605_12556
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle M2Retinexformer: Multi-Modal Retinexformer for Low-Light Image Enhancement
Aboelwafa, Youssef
Elmongui, Hicham G.
Torki, Marwan
Computer Vision and Pattern Recognition
Low-light image enhancement is challenging due to complex degradations, including amplified noise, artifacts, and color distortion. While Retinex-based deep learning methods have achieved promising results, they primarily rely on single-modality RGB information. We propose M2Retinexformer (Multi-Modal Retinexformer), a novel framework that extends Retinexformer by incorporating depth cues, luminance priors, and semantic features within a progressive refinement pipeline. Depth provides geometric context that is invariant to lighting variations, while luminance and semantic features offer explicit guidance on brightness distribution and scene understanding. Modalities are extracted at multiple scales and fused through cross-attention, with adaptive gating dynamically balancing illumination-guided self-attention and cross-attention based on the reliability of auxiliary cues. Evaluations on the LOL, SID, SMID, and SDSD benchmarks demonstrate overall improvements over Retinexformer and recent state-of-the-art methods. Code and pretrained weights are available at https://github.com/YoussefAboelwafa/M2Retinexformer
title M2Retinexformer: Multi-Modal Retinexformer for Low-Light Image Enhancement
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.12556