SMART-Editor: A Multi-Agent Framework for Human-Like Design Editing with Structural Integrity

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mondal, Ishani, Bharadwaj, Meera, Roy, Ayush, Garimella, Aparna, Boyd-Graber, Jordan Lee
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909722942111744
author Mondal, Ishani
Bharadwaj, Meera
Roy, Ayush
Garimella, Aparna
Boyd-Graber, Jordan Lee
author_facet Mondal, Ishani
Bharadwaj, Meera
Roy, Ayush
Garimella, Aparna
Boyd-Graber, Jordan Lee
contents We present SMART-Editor, a framework for compositional layout and content editing across structured (posters, websites) and unstructured (natural images) domains. Unlike prior models that perform local edits, SMART-Editor preserves global coherence through two strategies: Reward-Refine, an inference-time rewardguided refinement method, and RewardDPO, a training-time preference optimization approach using reward-aligned layout pairs. To evaluate model performance, we introduce SMARTEdit-Bench, a benchmark covering multi-domain, cascading edit scenarios. SMART-Editor outperforms strong baselines like InstructPix2Pix and HIVE, with RewardDPO achieving up to 15% gains in structured settings and Reward-Refine showing advantages on natural images. Automatic and human evaluations confirm the value of reward-guided planning in producing semantically consistent and visually aligned edits.
format Preprint
id arxiv_https___arxiv_org_abs_2507_23095
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SMART-Editor: A Multi-Agent Framework for Human-Like Design Editing with Structural Integrity
Mondal, Ishani
Bharadwaj, Meera
Roy, Ayush
Garimella, Aparna
Boyd-Graber, Jordan Lee
Computation and Language
Artificial Intelligence
We present SMART-Editor, a framework for compositional layout and content editing across structured (posters, websites) and unstructured (natural images) domains. Unlike prior models that perform local edits, SMART-Editor preserves global coherence through two strategies: Reward-Refine, an inference-time rewardguided refinement method, and RewardDPO, a training-time preference optimization approach using reward-aligned layout pairs. To evaluate model performance, we introduce SMARTEdit-Bench, a benchmark covering multi-domain, cascading edit scenarios. SMART-Editor outperforms strong baselines like InstructPix2Pix and HIVE, with RewardDPO achieving up to 15% gains in structured settings and Reward-Refine showing advantages on natural images. Automatic and human evaluations confirm the value of reward-guided planning in producing semantically consistent and visually aligned edits.
title SMART-Editor: A Multi-Agent Framework for Human-Like Design Editing with Structural Integrity
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2507.23095