SMART-Editor: A Multi-Agent Framework for Human-Like Design Editing with Structural Integrity
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909722942111744 |
|---|---|
| author | Mondal, Ishani Bharadwaj, Meera Roy, Ayush Garimella, Aparna Boyd-Graber, Jordan Lee |
| author_facet | Mondal, Ishani Bharadwaj, Meera Roy, Ayush Garimella, Aparna Boyd-Graber, Jordan Lee |
| contents | We present SMART-Editor, a framework for compositional layout and content editing across structured (posters, websites) and unstructured (natural images) domains. Unlike prior models that perform local edits, SMART-Editor preserves global coherence through two strategies: Reward-Refine, an inference-time rewardguided refinement method, and RewardDPO, a training-time preference optimization approach using reward-aligned layout pairs. To evaluate model performance, we introduce SMARTEdit-Bench, a benchmark covering multi-domain, cascading edit scenarios. SMART-Editor outperforms strong baselines like InstructPix2Pix and HIVE, with RewardDPO achieving up to 15% gains in structured settings and Reward-Refine showing advantages on natural images. Automatic and human evaluations confirm the value of reward-guided planning in producing semantically consistent and visually aligned edits. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2507_23095 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | SMART-Editor: A Multi-Agent Framework for Human-Like Design Editing with Structural Integrity Mondal, Ishani Bharadwaj, Meera Roy, Ayush Garimella, Aparna Boyd-Graber, Jordan Lee Computation and Language Artificial Intelligence We present SMART-Editor, a framework for compositional layout and content editing across structured (posters, websites) and unstructured (natural images) domains. Unlike prior models that perform local edits, SMART-Editor preserves global coherence through two strategies: Reward-Refine, an inference-time rewardguided refinement method, and RewardDPO, a training-time preference optimization approach using reward-aligned layout pairs. To evaluate model performance, we introduce SMARTEdit-Bench, a benchmark covering multi-domain, cascading edit scenarios. SMART-Editor outperforms strong baselines like InstructPix2Pix and HIVE, with RewardDPO achieving up to 15% gains in structured settings and Reward-Refine showing advantages on natural images. Automatic and human evaluations confirm the value of reward-guided planning in producing semantically consistent and visually aligned edits. |
| title | SMART-Editor: A Multi-Agent Framework for Human-Like Design Editing with Structural Integrity |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2507.23095 |