POEM: Interactive Prompt Optimization for Enhancing Multimodal Reasoning of Large Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: He, Jianben, Wang, Xingbo, Liu, Shiyi, Wu, Guande, Silva, Claudio, Qu, Huamin
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916415299125248
author He, Jianben
Wang, Xingbo
Liu, Shiyi
Wu, Guande
Silva, Claudio
Qu, Huamin
author_facet He, Jianben
Wang, Xingbo
Liu, Shiyi
Wu, Guande
Silva, Claudio
Qu, Huamin
contents Large language models (LLMs) have exhibited impressive abilities for multimodal content comprehension and reasoning with proper prompting in zero- or few-shot settings. Despite the proliferation of interactive systems developed to support prompt engineering for LLMs across various tasks, most have primarily focused on textual or visual inputs, thus neglecting the complex interplay between modalities within multimodal inputs. This oversight hinders the development of effective prompts that guide model multimodal reasoning processes by fully exploiting the rich context provided by multiple modalities. In this paper, we present POEM, a visual analytics system to facilitate efficient prompt engineering for enhancing the multimodal reasoning performance of LLMs. The system enables users to explore the interaction patterns across modalities at varying levels of detail for a comprehensive understanding of the multimodal knowledge elicited by various prompts. Through diverse recommendations of demonstration examples and instructional principles, POEM supports users in iteratively crafting and refining prompts to better align and enhance model knowledge with human insights. The effectiveness and efficiency of our system are validated through two case studies and interviews with experts.
format Preprint
id arxiv_https___arxiv_org_abs_2406_03843
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle POEM: Interactive Prompt Optimization for Enhancing Multimodal Reasoning of Large Language Models
He, Jianben
Wang, Xingbo
Liu, Shiyi
Wu, Guande
Silva, Claudio
Qu, Huamin
Human-Computer Interaction
Artificial Intelligence
68
H.5; I.2.1
Large language models (LLMs) have exhibited impressive abilities for multimodal content comprehension and reasoning with proper prompting in zero- or few-shot settings. Despite the proliferation of interactive systems developed to support prompt engineering for LLMs across various tasks, most have primarily focused on textual or visual inputs, thus neglecting the complex interplay between modalities within multimodal inputs. This oversight hinders the development of effective prompts that guide model multimodal reasoning processes by fully exploiting the rich context provided by multiple modalities. In this paper, we present POEM, a visual analytics system to facilitate efficient prompt engineering for enhancing the multimodal reasoning performance of LLMs. The system enables users to explore the interaction patterns across modalities at varying levels of detail for a comprehensive understanding of the multimodal knowledge elicited by various prompts. Through diverse recommendations of demonstration examples and instructional principles, POEM supports users in iteratively crafting and refining prompts to better align and enhance model knowledge with human insights. The effectiveness and efficiency of our system are validated through two case studies and interviews with experts.
title POEM: Interactive Prompt Optimization for Enhancing Multimodal Reasoning of Large Language Models
topic Human-Computer Interaction
Artificial Intelligence
68
H.5; I.2.1
url https://arxiv.org/abs/2406.03843