ReasonX: MLLM-Guided Intrinsic Image Decomposition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dirik, Alara, Wang, Tuanfeng, Ceylan, Duygu, Zafeiriou, Stefanos, Frühstück, Anna
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915653157388288
author Dirik, Alara
Wang, Tuanfeng
Ceylan, Duygu
Zafeiriou, Stefanos
Frühstück, Anna
author_facet Dirik, Alara
Wang, Tuanfeng
Ceylan, Duygu
Zafeiriou, Stefanos
Frühstück, Anna
contents Intrinsic image decomposition aims to separate images into physical components such as albedo, depth, normals, and illumination. While recent diffusion- and transformer-based models benefit from paired supervision from synthetic datasets, their generalization to diverse, real-world scenarios remains challenging. We propose ReasonX, a novel framework that leverages a multimodal large language model (MLLM) as a perceptual judge providing relative intrinsic comparisons, and uses these comparisons as GRPO rewards for fine-tuning intrinsic decomposition models on unlabeled, in-the-wild images. Unlike RL methods for generative models, our framework aligns conditional intrinsic predictors by rewarding agreement between the judge's relational assessments and analytically derived relations from the model's outputs. ReasonX is model-agnostic and can be applied to different intrinsic predictors. Across multiple base architectures and modalities, ReasonX yields significant improvements, including 9-25% WHDR reduction on IIW albedo and up to 46% depth accuracy gains on ETH3D, highlighting the promise of MLLM-guided comparative supervision to bridge low- and high-level vision reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2512_04222
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ReasonX: MLLM-Guided Intrinsic Image Decomposition
Dirik, Alara
Wang, Tuanfeng
Ceylan, Duygu
Zafeiriou, Stefanos
Frühstück, Anna
Computer Vision and Pattern Recognition
Intrinsic image decomposition aims to separate images into physical components such as albedo, depth, normals, and illumination. While recent diffusion- and transformer-based models benefit from paired supervision from synthetic datasets, their generalization to diverse, real-world scenarios remains challenging. We propose ReasonX, a novel framework that leverages a multimodal large language model (MLLM) as a perceptual judge providing relative intrinsic comparisons, and uses these comparisons as GRPO rewards for fine-tuning intrinsic decomposition models on unlabeled, in-the-wild images. Unlike RL methods for generative models, our framework aligns conditional intrinsic predictors by rewarding agreement between the judge's relational assessments and analytically derived relations from the model's outputs. ReasonX is model-agnostic and can be applied to different intrinsic predictors. Across multiple base architectures and modalities, ReasonX yields significant improvements, including 9-25% WHDR reduction on IIW albedo and up to 46% depth accuracy gains on ETH3D, highlighting the promise of MLLM-guided comparative supervision to bridge low- and high-level vision reasoning.
title ReasonX: MLLM-Guided Intrinsic Image Decomposition
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.04222