MiLDEdit: Reasoning-Based Multi-Layer Design Document Editing

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lin, Zihao, Zhu, Wanrong, Gu, Jiuxiang, Kil, Jihyung, Tensmeyer, Christopher, Zhang, Lin, Liu, Shilong, Zhang, Ruiyi, Huang, Lifu, Morariu, Vlad I., Sun, Tong
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908795403239424
author Lin, Zihao
Zhu, Wanrong
Gu, Jiuxiang
Kil, Jihyung
Tensmeyer, Christopher
Zhang, Lin
Liu, Shilong
Zhang, Ruiyi
Huang, Lifu
Morariu, Vlad I.
Sun, Tong
author_facet Lin, Zihao
Zhu, Wanrong
Gu, Jiuxiang
Kil, Jihyung
Tensmeyer, Christopher
Zhang, Lin
Liu, Shilong
Zhang, Ruiyi
Huang, Lifu
Morariu, Vlad I.
Sun, Tong
contents Real-world design documents (e.g., posters) are inherently multi-layered, combining decoration, text, and images. Editing them from natural-language instructions requires fine-grained, layer-aware reasoning to identify relevant layers and coordinate modifications. Prior work largely overlooks multi-layer design document editing, focusing instead on single-layer image editing or multi-layer generation, which assume a flat canvas and lack the reasoning needed to determine what and where to modify. To address this gap, we introduce the Multi-Layer Document Editing Agent (MiLDEAgent), a reasoning-based framework that combines an RL-trained multimodal reasoner for layer-wise understanding with an image editor for targeted modifications. To systematically benchmark this setting, we introduce the MiLDEBench, a human-in-the-loop corpus of over 20K design documents paired with diverse editing instructions. The benchmark is complemented by a task-specific evaluation protocol, MiLDEEval, which spans four dimensions including instruction following, layout consistency, aesthetics, and text rendering. Extensive experiments on 14 open-source and 2 closed-source models reveal that existing approaches fail to generalize: open-source models often cannot complete multi-layer document editing tasks, while closed-source models suffer from format violations. In contrast, MiLDEAgent achieves strong layer-aware reasoning and precise editing, significantly outperforming all open-source baselines and attaining performance comparable to closed-source models, thereby establishing the first strong baseline for multi-layer document editing.
format Preprint
id arxiv_https___arxiv_org_abs_2601_04589
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MiLDEdit: Reasoning-Based Multi-Layer Design Document Editing
Lin, Zihao
Zhu, Wanrong
Gu, Jiuxiang
Kil, Jihyung
Tensmeyer, Christopher
Zhang, Lin
Liu, Shilong
Zhang, Ruiyi
Huang, Lifu
Morariu, Vlad I.
Sun, Tong
Computer Vision and Pattern Recognition
Real-world design documents (e.g., posters) are inherently multi-layered, combining decoration, text, and images. Editing them from natural-language instructions requires fine-grained, layer-aware reasoning to identify relevant layers and coordinate modifications. Prior work largely overlooks multi-layer design document editing, focusing instead on single-layer image editing or multi-layer generation, which assume a flat canvas and lack the reasoning needed to determine what and where to modify. To address this gap, we introduce the Multi-Layer Document Editing Agent (MiLDEAgent), a reasoning-based framework that combines an RL-trained multimodal reasoner for layer-wise understanding with an image editor for targeted modifications. To systematically benchmark this setting, we introduce the MiLDEBench, a human-in-the-loop corpus of over 20K design documents paired with diverse editing instructions. The benchmark is complemented by a task-specific evaluation protocol, MiLDEEval, which spans four dimensions including instruction following, layout consistency, aesthetics, and text rendering. Extensive experiments on 14 open-source and 2 closed-source models reveal that existing approaches fail to generalize: open-source models often cannot complete multi-layer document editing tasks, while closed-source models suffer from format violations. In contrast, MiLDEAgent achieves strong layer-aware reasoning and precise editing, significantly outperforming all open-source baselines and attaining performance comparable to closed-source models, thereby establishing the first strong baseline for multi-layer document editing.
title MiLDEdit: Reasoning-Based Multi-Layer Design Document Editing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.04589