SmartPhotoCrafter: Unified Reasoning, Generation and Optimization for Automatic Photographic Image Editing

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zeng, Ying, Luo, Miaosen, Li, Guangyuan, Yang, Yang, Fan, Ruiyang, Shi, Linxiao, Yang, Qirui, Zhang, Jian, Liu, Chengcheng, Zheng, Siming, Chen, Jinwei, Li, Bo, Jiang, Peng-Tao
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918459728723968
author Zeng, Ying
Luo, Miaosen
Li, Guangyuan
Yang, Yang
Fan, Ruiyang
Shi, Linxiao
Yang, Qirui
Zhang, Jian
Liu, Chengcheng
Zheng, Siming
Chen, Jinwei
Li, Bo
Jiang, Peng-Tao
author_facet Zeng, Ying
Luo, Miaosen
Li, Guangyuan
Yang, Yang
Fan, Ruiyang
Shi, Linxiao
Yang, Qirui
Zhang, Jian
Liu, Chengcheng
Zheng, Siming
Chen, Jinwei
Li, Bo
Jiang, Peng-Tao
contents Traditional photographic image editing typically requires users to possess sufficient aesthetic understanding to provide appropriate instructions for adjusting image quality and camera parameters. However, this paradigm relies on explicit human instruction of aesthetic intent, which is often ambiguous, incomplete, or inaccessible to non-expert users. In this work, we propose SmartPhotoCrafter, an automatic photographic image editing method which formulates image editing as a tightly coupled reasoning-to-generation process. The proposed model first performs image quality comprehension and identifies deficiencies by the Image Critic module, and then the Photographic Artist module realizes targeted edits to enhance image appeal, eliminating the need for explicit human instructions. A multi-stage training pipeline is adopted: (i) Foundation pretraining to establish basic aesthetic understanding and editing capabilities, (ii) Adaptation with reasoning-guided multi-edit supervision to incorporate rich semantic guidance, and (iii) Coordinated reasoning-to generation reinforcement learning to jointly optimize reasoning and generation. During training, SmartPhotoCrafter emphasizes photo-realistic image generation, while supporting both image restoration and retouching tasks with consistent adherence to color- and tone-related semantics. We also construct a stage-specific dataset, which progressively builds reasoning and controllable generation, effective cross-module collaboration, and ultimately high-quality photographic enhancement. Experiments demonstrate that SmartPhotoCrafter outperforms existing generative models on the task of automatic photographic enhancement, achieving photo-realistic results while exhibiting higher tonal sensitivity to retouching instructions. Project page: https://github.com/vivoCameraResearch/SmartPhotoCrafter.
format Preprint
id arxiv_https___arxiv_org_abs_2604_19587
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SmartPhotoCrafter: Unified Reasoning, Generation and Optimization for Automatic Photographic Image Editing
Zeng, Ying
Luo, Miaosen
Li, Guangyuan
Yang, Yang
Fan, Ruiyang
Shi, Linxiao
Yang, Qirui
Zhang, Jian
Liu, Chengcheng
Zheng, Siming
Chen, Jinwei
Li, Bo
Jiang, Peng-Tao
Computer Vision and Pattern Recognition
Traditional photographic image editing typically requires users to possess sufficient aesthetic understanding to provide appropriate instructions for adjusting image quality and camera parameters. However, this paradigm relies on explicit human instruction of aesthetic intent, which is often ambiguous, incomplete, or inaccessible to non-expert users. In this work, we propose SmartPhotoCrafter, an automatic photographic image editing method which formulates image editing as a tightly coupled reasoning-to-generation process. The proposed model first performs image quality comprehension and identifies deficiencies by the Image Critic module, and then the Photographic Artist module realizes targeted edits to enhance image appeal, eliminating the need for explicit human instructions. A multi-stage training pipeline is adopted: (i) Foundation pretraining to establish basic aesthetic understanding and editing capabilities, (ii) Adaptation with reasoning-guided multi-edit supervision to incorporate rich semantic guidance, and (iii) Coordinated reasoning-to generation reinforcement learning to jointly optimize reasoning and generation. During training, SmartPhotoCrafter emphasizes photo-realistic image generation, while supporting both image restoration and retouching tasks with consistent adherence to color- and tone-related semantics. We also construct a stage-specific dataset, which progressively builds reasoning and controllable generation, effective cross-module collaboration, and ultimately high-quality photographic enhancement. Experiments demonstrate that SmartPhotoCrafter outperforms existing generative models on the task of automatic photographic enhancement, achieving photo-realistic results while exhibiting higher tonal sensitivity to retouching instructions. Project page: https://github.com/vivoCameraResearch/SmartPhotoCrafter.
title SmartPhotoCrafter: Unified Reasoning, Generation and Optimization for Automatic Photographic Image Editing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.19587