Visualizing Dialogues: Enhancing Image Selection through Dialogue Understanding with Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kao, Chang-Sheng, Chen, Yun-Nung
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929409055784960
author Kao, Chang-Sheng
Chen, Yun-Nung
author_facet Kao, Chang-Sheng
Chen, Yun-Nung
contents Recent advancements in dialogue systems have highlighted the significance of integrating multimodal responses, which enable conveying ideas through diverse modalities rather than solely relying on text-based interactions. This enrichment not only improves overall communicative efficacy but also enhances the quality of conversational experiences. However, existing methods for dialogue-to-image retrieval face limitations due to the constraints of pre-trained vision language models (VLMs) in comprehending complex dialogues accurately. To address this, we present a novel approach leveraging the robust reasoning capabilities of large language models (LLMs) to generate precise dialogue-associated visual descriptors, facilitating seamless connection with images. Extensive experiments conducted on benchmark data validate the effectiveness of our proposed approach in deriving concise and accurate visual descriptors, leading to significant enhancements in dialogue-to-image retrieval performance. Furthermore, our findings demonstrate the method's generalizability across diverse visual cues, various LLMs, and different datasets, underscoring its practicality and potential impact in real-world applications.
format Preprint
id arxiv_https___arxiv_org_abs_2407_03615
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Visualizing Dialogues: Enhancing Image Selection through Dialogue Understanding with Large Language Models
Kao, Chang-Sheng
Chen, Yun-Nung
Computation and Language
Recent advancements in dialogue systems have highlighted the significance of integrating multimodal responses, which enable conveying ideas through diverse modalities rather than solely relying on text-based interactions. This enrichment not only improves overall communicative efficacy but also enhances the quality of conversational experiences. However, existing methods for dialogue-to-image retrieval face limitations due to the constraints of pre-trained vision language models (VLMs) in comprehending complex dialogues accurately. To address this, we present a novel approach leveraging the robust reasoning capabilities of large language models (LLMs) to generate precise dialogue-associated visual descriptors, facilitating seamless connection with images. Extensive experiments conducted on benchmark data validate the effectiveness of our proposed approach in deriving concise and accurate visual descriptors, leading to significant enhancements in dialogue-to-image retrieval performance. Furthermore, our findings demonstrate the method's generalizability across diverse visual cues, various LLMs, and different datasets, underscoring its practicality and potential impact in real-world applications.
title Visualizing Dialogues: Enhancing Image Selection through Dialogue Understanding with Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2407.03615