AceTone: Bridging Words and Colors for Conditional Image Grading

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Ma, Tianren, Liao, Mingxiang, Zhang, Xijin, Ye, Qixiang
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910093105168384
author Ma, Tianren
Liao, Mingxiang
Zhang, Xijin
Ye, Qixiang
author_facet Ma, Tianren
Liao, Mingxiang
Zhang, Xijin
Ye, Qixiang
contents Color affects how we interpret image style and emotion. Previous color grading methods rely on patch-wise recoloring or fixed filter banks, struggling to generalize across creative intents or align with human aesthetic preferences. In this study, we propose AceTone, the first approach that supports multimodal conditioned color grading within a unified framework. AceTone formulates grading as a generative color transformation task, where a model directly produces 3D-LUTs conditioned on text prompts or reference images. We develop a VQ-VAE based tokenizer which compresses a $3\times32^3$ LUT vector to 64 discrete tokens with $ΔE<2$ fidelity. We further build a large-scale dataset, AceTone-800K, and train a vision-language model to predict LUT tokens, followed by reinforcement learning to align outputs with perceptual fidelity and aesthetics. Experiments show that AceTone achieves state-of-the-art performance on both text-guided and reference-guided grading tasks, improving LPIPS by up to 50% over existing methods. Human evaluations confirm that AceTone's results are visually pleasing and stylistically coherent, demonstrating a new pathway toward language-driven, aesthetic-aligned color grading.
format Preprint
id arxiv_https___arxiv_org_abs_2604_00530
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AceTone: Bridging Words and Colors for Conditional Image Grading
Ma, Tianren
Liao, Mingxiang
Zhang, Xijin
Ye, Qixiang
Computer Vision and Pattern Recognition
Color affects how we interpret image style and emotion. Previous color grading methods rely on patch-wise recoloring or fixed filter banks, struggling to generalize across creative intents or align with human aesthetic preferences. In this study, we propose AceTone, the first approach that supports multimodal conditioned color grading within a unified framework. AceTone formulates grading as a generative color transformation task, where a model directly produces 3D-LUTs conditioned on text prompts or reference images. We develop a VQ-VAE based tokenizer which compresses a $3\times32^3$ LUT vector to 64 discrete tokens with $ΔE<2$ fidelity. We further build a large-scale dataset, AceTone-800K, and train a vision-language model to predict LUT tokens, followed by reinforcement learning to align outputs with perceptual fidelity and aesthetics. Experiments show that AceTone achieves state-of-the-art performance on both text-guided and reference-guided grading tasks, improving LPIPS by up to 50% over existing methods. Human evaluations confirm that AceTone's results are visually pleasing and stylistically coherent, demonstrating a new pathway toward language-driven, aesthetic-aligned color grading.
title AceTone: Bridging Words and Colors for Conditional Image Grading
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.00530