DesignProbe: A Graphic Design Benchmark for Multimodal Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lin, Jieru, Huang, Danqing, Zhao, Tiejun, Zhan, Dechen, Lin, Chin-Yew
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910418770853888
author Lin, Jieru
Huang, Danqing
Zhao, Tiejun
Zhan, Dechen
Lin, Chin-Yew
author_facet Lin, Jieru
Huang, Danqing
Zhao, Tiejun
Zhan, Dechen
Lin, Chin-Yew
contents A well-executed graphic design typically achieves harmony in two levels, from the fine-grained design elements (color, font and layout) to the overall design. This complexity makes the comprehension of graphic design challenging, for it needs the capability to both recognize the design elements and understand the design. With the rapid development of Multimodal Large Language Models (MLLMs), we establish the DesignProbe, a benchmark to investigate the capability of MLLMs in design. Our benchmark includes eight tasks in total, across both the fine-grained element level and the overall design level. At design element level, we consider both the attribute recognition and semantic understanding tasks. At overall design level, we include style and metaphor. 9 MLLMs are tested and we apply GPT-4 as evaluator. Besides, further experiments indicates that refining prompts can enhance the performance of MLLMs. We first rewrite the prompts by different LLMs and found increased performances appear in those who self-refined by their own LLMs. We then add extra task knowledge in two different ways (text descriptions and image examples), finding that adding images boost much more performance over texts.
format Preprint
id arxiv_https___arxiv_org_abs_2404_14801
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DesignProbe: A Graphic Design Benchmark for Multimodal Large Language Models
Lin, Jieru
Huang, Danqing
Zhao, Tiejun
Zhan, Dechen
Lin, Chin-Yew
Computer Vision and Pattern Recognition
A well-executed graphic design typically achieves harmony in two levels, from the fine-grained design elements (color, font and layout) to the overall design. This complexity makes the comprehension of graphic design challenging, for it needs the capability to both recognize the design elements and understand the design. With the rapid development of Multimodal Large Language Models (MLLMs), we establish the DesignProbe, a benchmark to investigate the capability of MLLMs in design. Our benchmark includes eight tasks in total, across both the fine-grained element level and the overall design level. At design element level, we consider both the attribute recognition and semantic understanding tasks. At overall design level, we include style and metaphor. 9 MLLMs are tested and we apply GPT-4 as evaluator. Besides, further experiments indicates that refining prompts can enhance the performance of MLLMs. We first rewrite the prompts by different LLMs and found increased performances appear in those who self-refined by their own LLMs. We then add extra task knowledge in two different ways (text descriptions and image examples), finding that adding images boost much more performance over texts.
title DesignProbe: A Graphic Design Benchmark for Multimodal Large Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2404.14801