Saved in:
Bibliographic Details
Main Authors: Cheng, Jiaxin, Zhao, Zixu, He, Tong, Xiao, Tianjun, Zhou, Yicong, Zhang, Zheng
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2409.04847
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915098141917184
author Cheng, Jiaxin
Zhao, Zixu
He, Tong
Xiao, Tianjun
Zhou, Yicong
Zhang, Zheng
author_facet Cheng, Jiaxin
Zhao, Zixu
He, Tong
Xiao, Tianjun
Zhou, Yicong
Zhang, Zheng
contents Recent advancements in generative models have significantly enhanced their capacity for image generation, enabling a wide range of applications such as image editing, completion and video editing. A specialized area within generative modeling is layout-to-image (L2I) generation, where predefined layouts of objects guide the generative process. In this study, we introduce a novel regional cross-attention module tailored to enrich layout-to-image generation. This module notably improves the representation of layout regions, particularly in scenarios where existing methods struggle with highly complex and detailed textual descriptions. Moreover, while current open-vocabulary L2I methods are trained in an open-set setting, their evaluations often occur in closed-set environments. To bridge this gap, we propose two metrics to assess L2I performance in open-vocabulary scenarios. Additionally, we conduct a comprehensive user study to validate the consistency of these metrics with human preferences.
format Preprint
id arxiv_https___arxiv_org_abs_2409_04847
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Rethinking The Training And Evaluation of Rich-Context Layout-to-Image Generation
Cheng, Jiaxin
Zhao, Zixu
He, Tong
Xiao, Tianjun
Zhou, Yicong
Zhang, Zheng
Computer Vision and Pattern Recognition
Recent advancements in generative models have significantly enhanced their capacity for image generation, enabling a wide range of applications such as image editing, completion and video editing. A specialized area within generative modeling is layout-to-image (L2I) generation, where predefined layouts of objects guide the generative process. In this study, we introduce a novel regional cross-attention module tailored to enrich layout-to-image generation. This module notably improves the representation of layout regions, particularly in scenarios where existing methods struggle with highly complex and detailed textual descriptions. Moreover, while current open-vocabulary L2I methods are trained in an open-set setting, their evaluations often occur in closed-set environments. To bridge this gap, we propose two metrics to assess L2I performance in open-vocabulary scenarios. Additionally, we conduct a comprehensive user study to validate the consistency of these metrics with human preferences.
title Rethinking The Training And Evaluation of Rich-Context Layout-to-Image Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2409.04847