MiLDEdit: Reasoning-Based Multi-Layer Design Document Editing
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866908795403239424 |
|---|---|
| author | Lin, Zihao Zhu, Wanrong Gu, Jiuxiang Kil, Jihyung Tensmeyer, Christopher Zhang, Lin Liu, Shilong Zhang, Ruiyi Huang, Lifu Morariu, Vlad I. Sun, Tong |
| author_facet | Lin, Zihao Zhu, Wanrong Gu, Jiuxiang Kil, Jihyung Tensmeyer, Christopher Zhang, Lin Liu, Shilong Zhang, Ruiyi Huang, Lifu Morariu, Vlad I. Sun, Tong |
| contents | Real-world design documents (e.g., posters) are inherently multi-layered, combining decoration, text, and images. Editing them from natural-language instructions requires fine-grained, layer-aware reasoning to identify relevant layers and coordinate modifications. Prior work largely overlooks multi-layer design document editing, focusing instead on single-layer image editing or multi-layer generation, which assume a flat canvas and lack the reasoning needed to determine what and where to modify. To address this gap, we introduce the Multi-Layer Document Editing Agent (MiLDEAgent), a reasoning-based framework that combines an RL-trained multimodal reasoner for layer-wise understanding with an image editor for targeted modifications. To systematically benchmark this setting, we introduce the MiLDEBench, a human-in-the-loop corpus of over 20K design documents paired with diverse editing instructions. The benchmark is complemented by a task-specific evaluation protocol, MiLDEEval, which spans four dimensions including instruction following, layout consistency, aesthetics, and text rendering. Extensive experiments on 14 open-source and 2 closed-source models reveal that existing approaches fail to generalize: open-source models often cannot complete multi-layer document editing tasks, while closed-source models suffer from format violations. In contrast, MiLDEAgent achieves strong layer-aware reasoning and precise editing, significantly outperforming all open-source baselines and attaining performance comparable to closed-source models, thereby establishing the first strong baseline for multi-layer document editing. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_04589 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | MiLDEdit: Reasoning-Based Multi-Layer Design Document Editing Lin, Zihao Zhu, Wanrong Gu, Jiuxiang Kil, Jihyung Tensmeyer, Christopher Zhang, Lin Liu, Shilong Zhang, Ruiyi Huang, Lifu Morariu, Vlad I. Sun, Tong Computer Vision and Pattern Recognition Real-world design documents (e.g., posters) are inherently multi-layered, combining decoration, text, and images. Editing them from natural-language instructions requires fine-grained, layer-aware reasoning to identify relevant layers and coordinate modifications. Prior work largely overlooks multi-layer design document editing, focusing instead on single-layer image editing or multi-layer generation, which assume a flat canvas and lack the reasoning needed to determine what and where to modify. To address this gap, we introduce the Multi-Layer Document Editing Agent (MiLDEAgent), a reasoning-based framework that combines an RL-trained multimodal reasoner for layer-wise understanding with an image editor for targeted modifications. To systematically benchmark this setting, we introduce the MiLDEBench, a human-in-the-loop corpus of over 20K design documents paired with diverse editing instructions. The benchmark is complemented by a task-specific evaluation protocol, MiLDEEval, which spans four dimensions including instruction following, layout consistency, aesthetics, and text rendering. Extensive experiments on 14 open-source and 2 closed-source models reveal that existing approaches fail to generalize: open-source models often cannot complete multi-layer document editing tasks, while closed-source models suffer from format violations. In contrast, MiLDEAgent achieves strong layer-aware reasoning and precise editing, significantly outperforming all open-source baselines and attaining performance comparable to closed-source models, thereby establishing the first strong baseline for multi-layer document editing. |
| title | MiLDEdit: Reasoning-Based Multi-Layer Design Document Editing |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2601.04589 |