CAL-RAG: Retrieval-Augmented Multi-Agent Generation for Content-Aware Layout Design

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Forouzandehmehr, Najmeh, Maragheh, Reza Yousefi, Kollipara, Sriram, Zhao, Kai, Biswas, Topojoy, Korpeoglu, Evren, Achan, Kannan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912453404655616
author Forouzandehmehr, Najmeh
Maragheh, Reza Yousefi
Kollipara, Sriram
Zhao, Kai
Biswas, Topojoy
Korpeoglu, Evren
Achan, Kannan
author_facet Forouzandehmehr, Najmeh
Maragheh, Reza Yousefi
Kollipara, Sriram
Zhao, Kai
Biswas, Topojoy
Korpeoglu, Evren
Achan, Kannan
contents Automated content-aware layout generation -- the task of arranging visual elements such as text, logos, and underlays on a background canvas -- remains a fundamental yet under-explored problem in intelligent design systems. While recent advances in deep generative models and large language models (LLMs) have shown promise in structured content generation, most existing approaches lack grounding in contextual design exemplars and fall short in handling semantic alignment and visual coherence. In this work we introduce CAL-RAG, a retrieval-augmented, agentic framework for content-aware layout generation that integrates multimodal retrieval, large language models, and collaborative agentic reasoning. Our system retrieves relevant layout examples from a structured knowledge base and invokes an LLM-based layout recommender to propose structured element placements. A vision-language grader agent evaluates the layout with visual metrics, and a feedback agent provides targeted refinements, enabling iterative improvement. We implement our framework using LangGraph and evaluate it on the PKU PosterLayout dataset, a benchmark rich in semantic and structural variability. CAL-RAG achieves state-of-the-art performance across multiple layout metrics -- including underlay effectiveness, element alignment, and overlap -- substantially outperforming strong baselines such as LayoutPrompter. These results demonstrate that combining retrieval augmentation with agentic multi-step reasoning yields a scalable, interpretable, and high-fidelity solution for automated layout generation.
format Preprint
id arxiv_https___arxiv_org_abs_2506_21934
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CAL-RAG: Retrieval-Augmented Multi-Agent Generation for Content-Aware Layout Design
Forouzandehmehr, Najmeh
Maragheh, Reza Yousefi
Kollipara, Sriram
Zhao, Kai
Biswas, Topojoy
Korpeoglu, Evren
Achan, Kannan
Information Retrieval
Computer Vision and Pattern Recognition
I.3.3; I.2.11; H.5.2
Automated content-aware layout generation -- the task of arranging visual elements such as text, logos, and underlays on a background canvas -- remains a fundamental yet under-explored problem in intelligent design systems. While recent advances in deep generative models and large language models (LLMs) have shown promise in structured content generation, most existing approaches lack grounding in contextual design exemplars and fall short in handling semantic alignment and visual coherence. In this work we introduce CAL-RAG, a retrieval-augmented, agentic framework for content-aware layout generation that integrates multimodal retrieval, large language models, and collaborative agentic reasoning. Our system retrieves relevant layout examples from a structured knowledge base and invokes an LLM-based layout recommender to propose structured element placements. A vision-language grader agent evaluates the layout with visual metrics, and a feedback agent provides targeted refinements, enabling iterative improvement. We implement our framework using LangGraph and evaluate it on the PKU PosterLayout dataset, a benchmark rich in semantic and structural variability. CAL-RAG achieves state-of-the-art performance across multiple layout metrics -- including underlay effectiveness, element alignment, and overlap -- substantially outperforming strong baselines such as LayoutPrompter. These results demonstrate that combining retrieval augmentation with agentic multi-step reasoning yields a scalable, interpretable, and high-fidelity solution for automated layout generation.
title CAL-RAG: Retrieval-Augmented Multi-Agent Generation for Content-Aware Layout Design
topic Information Retrieval
Computer Vision and Pattern Recognition
I.3.3; I.2.11; H.5.2
url https://arxiv.org/abs/2506.21934