RetouchLLM: Training-free Code-based Image Retouching with Vision Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Ye-Bin, Moon, Miles, Roy, Oh, Tae-Hyun, Elezi, Ismail, Deng, Jiankang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
$V_kD:$ Improving Knowledge Distillation using Orthogonal Projections
by: Miles, Roy, et al.
Published: (2024)
by: Miles, Roy, et al.
Published: (2024)
VeLoRA: Memory Efficient Training using Rank-1 Sub-Token Projections
by: Miles, Roy, et al.
Published: (2024)
by: Miles, Roy, et al.
Published: (2024)
DiffRetouch: Using Diffusion to Retouch on the Shoulder of Experts
by: Duan, Zheng-Peng, et al.
Published: (2024)
by: Duan, Zheng-Peng, et al.
Published: (2024)
RetouchIQ: MLLM Agents for Instruction-Based Image Retouching with Generalist Reward
by: Wu, Qiucheng, et al.
Published: (2026)
by: Wu, Qiucheng, et al.
Published: (2026)
Agentic Retoucher for Text-To-Image Generation
by: Shen, Shaocheng, et al.
Published: (2026)
by: Shen, Shaocheng, et al.
Published: (2026)
G3DR: Generative 3D Reconstruction in ImageNet
by: Reddy, Pradyumna, et al.
Published: (2024)
by: Reddy, Pradyumna, et al.
Published: (2024)
VeraRetouch: A Lightweight Fully Differentiable Framework for Multi-Task Reasoning Photo Retouching
by: Guo, Yihong, et al.
Published: (2026)
by: Guo, Yihong, et al.
Published: (2026)
SATGround: A Spatially-Aware Approach for Visual Grounding in Remote Sensing
by: Toker, Aysim, et al.
Published: (2025)
by: Toker, Aysim, et al.
Published: (2025)
Label-guided Facial Retouching Reversion
by: Zhao, Guanhua, et al.
Published: (2024)
by: Zhao, Guanhua, et al.
Published: (2024)
"Principal Components" Enable A New Language of Images
by: Wen, Xin, et al.
Published: (2025)
by: Wen, Xin, et al.
Published: (2025)
Taming Lookup Tables for Efficient Image Retouching
by: Yang, Sidi, et al.
Published: (2024)
by: Yang, Sidi, et al.
Published: (2024)
Type-R: Automatically Retouching Typos for Text-to-Image Generation
by: Shimoda, Wataru, et al.
Published: (2024)
by: Shimoda, Wataru, et al.
Published: (2024)
Do You See What I Am Pointing At? Gesture-Based Egocentric Video Question Answering
by: Choi, Yura, et al.
Published: (2026)
by: Choi, Yura, et al.
Published: (2026)
Deep Active Learning: A Reality Check
by: Gashi, Edrina, et al.
Published: (2024)
by: Gashi, Edrina, et al.
Published: (2024)
MoFRR: Mixture of Diffusion Models for Face Retouching Restoration
by: Liu, Jiaxin, et al.
Published: (2025)
by: Liu, Jiaxin, et al.
Published: (2025)
PerTouch: VLM-Driven Agent for Personalized and Semantic Image Retouching
by: Chang, Zewei, et al.
Published: (2025)
by: Chang, Zewei, et al.
Published: (2025)
Content-Adaptive Image Retouching Guided by Attribute-Based Text Representation
by: Zhu, Hancheng, et al.
Published: (2025)
by: Zhu, Hancheng, et al.
Published: (2025)
PhotoArtAgent: Intelligent Photo Retouching with Language Model-Based Artist Agents
by: Chen, Haoyu, et al.
Published: (2025)
by: Chen, Haoyu, et al.
Published: (2025)
Controllable and Gradual Facial Blemishes Retouching via Physics-Based Modelling
by: Shuai, Chenhao, et al.
Published: (2024)
by: Shuai, Chenhao, et al.
Published: (2024)
FRRffusion: Unveiling Authenticity with Diffusion-Based Face Retouching Reversal
by: Xing, Fengchuang, et al.
Published: (2024)
by: Xing, Fengchuang, et al.
Published: (2024)
Detection of Digital Facial Retouching utilizing Face Beauty Information
by: Srock, Philipp, et al.
Published: (2025)
by: Srock, Philipp, et al.
Published: (2025)
MonetGPT: Solving Puzzles Enhances MLLMs' Image Retouching Skills
by: Dutt, Niladri Shekhar, et al.
Published: (2025)
by: Dutt, Niladri Shekhar, et al.
Published: (2025)
NCST: Neural-based Color Style Transfer for Video Retouching
by: Jiang, Xintao, et al.
Published: (2024)
by: Jiang, Xintao, et al.
Published: (2024)
VLM's Eye Examination: Instruct and Inspect Visual Competency of Vision Language Models
by: Hyeon-Woo, Nam, et al.
Published: (2024)
by: Hyeon-Woo, Nam, et al.
Published: (2024)
CameraMaster: Unified Camera Semantic-Parameter Control for Photography Retouching
by: Yang, Qirui, et al.
Published: (2025)
by: Yang, Qirui, et al.
Published: (2025)
INRetouch: Context Aware Implicit Neural Representation for Photography Retouching
by: Elezabi, Omar, et al.
Published: (2024)
by: Elezabi, Omar, et al.
Published: (2024)
Layout-and-Retouch: A Dual-stage Framework for Improving Diversity in Personalized Image Generation
by: Kim, Kangyeol, et al.
Published: (2024)
by: Kim, Kangyeol, et al.
Published: (2024)
Three Heads Are Better Than One: Complementary Experts for Long-Tailed Semi-supervised Learning
by: Ma, Chengcheng, et al.
Published: (2023)
by: Ma, Chengcheng, et al.
Published: (2023)
JarvisArt: Liberating Human Artistic Creativity via an Intelligent Photo Retouching Agent
by: Lin, Yunlong, et al.
Published: (2025)
by: Lin, Yunlong, et al.
Published: (2025)
BEAF: Observing BEfore-AFter Changes to Evaluate Hallucination in Vision-language Models
by: Ye-Bin, Moon, et al.
Published: (2024)
by: Ye-Bin, Moon, et al.
Published: (2024)
Fractal Calibration for long-tailed object detection
by: Alexandridis, Konstantinos Panagiotis, et al.
Published: (2024)
by: Alexandridis, Konstantinos Panagiotis, et al.
Published: (2024)
Region-based Cluster Discrimination for Visual Representation Learning
by: Xie, Yin, et al.
Published: (2025)
by: Xie, Yin, et al.
Published: (2025)
Draw Like an Artist: Complex Scene Generation with Diffusion Model via Composition, Painting, and Retouching
by: Liu, Minghao, et al.
Published: (2024)
by: Liu, Minghao, et al.
Published: (2024)
ViCToR: Improving Visual Comprehension via Token Reconstruction for Pretraining LMMs
by: Xie, Yin, et al.
Published: (2024)
by: Xie, Yin, et al.
Published: (2024)
BeautyGRPO: Aesthetic Alignment for Face Retouching via Dynamic Path Guidance and Fine-Grained Preference Modeling
by: Yang, Jiachen, et al.
Published: (2026)
by: Yang, Jiachen, et al.
Published: (2026)
Early Failure Detection and Intervention in Video Diffusion Models
by: Byung-Ki, Kwon, et al.
Published: (2026)
by: Byung-Ki, Kwon, et al.
Published: (2026)
Aligning What Vision-Language Models See and Perceive with Adaptive Information Flow
by: Liu, Chengxin, et al.
Published: (2026)
by: Liu, Chengxin, et al.
Published: (2026)
AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models
by: Sung-Bin, Kim, et al.
Published: (2024)
by: Sung-Bin, Kim, et al.
Published: (2024)
MemBench: Memorized Image Trigger Prompt Dataset for Diffusion Models
by: Hong, Chunsan, et al.
Published: (2024)
by: Hong, Chunsan, et al.
Published: (2024)
DreamCAD: Scaling Multi-modal CAD Generation using Differentiable Parametric Surfaces
by: Khan, Mohammad Sadil, et al.
Published: (2026)
by: Khan, Mohammad Sadil, et al.
Published: (2026)
Similar Items
-
$V_kD:$ Improving Knowledge Distillation using Orthogonal Projections
by: Miles, Roy, et al.
Published: (2024) -
VeLoRA: Memory Efficient Training using Rank-1 Sub-Token Projections
by: Miles, Roy, et al.
Published: (2024) -
DiffRetouch: Using Diffusion to Retouch on the Shoulder of Experts
by: Duan, Zheng-Peng, et al.
Published: (2024) -
RetouchIQ: MLLM Agents for Instruction-Based Image Retouching with Generalist Reward
by: Wu, Qiucheng, et al.
Published: (2026) -
Agentic Retoucher for Text-To-Image Generation
by: Shen, Shaocheng, et al.
Published: (2026)