Do Images Clarify? A Study on the Effect of Images on Clarifying Questions in Conversational Search

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Siro, Clemencia, Abbasiantaeb, Zahra, Yuan, Yifei, Aliannejadi, Mohammad, de Rijke, Maarten
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908822807773184
author Siro, Clemencia
Abbasiantaeb, Zahra
Yuan, Yifei
Aliannejadi, Mohammad
de Rijke, Maarten
author_facet Siro, Clemencia
Abbasiantaeb, Zahra
Yuan, Yifei
Aliannejadi, Mohammad
de Rijke, Maarten
contents Conversational search systems increasingly employ clarifying questions to refine user queries and improve the search experience. Previous studies have demonstrated the usefulness of text-based clarifying questions in enhancing both retrieval performance and user experience. While images have been shown to improve retrieval performance in various contexts, their impact on user performance when incorporated into clarifying questions remains largely unexplored. We conduct a user study with 73 participants to investigate the role of images in conversational search, specifically examining their effects on two search-related tasks: (i) answering clarifying questions and (ii) query reformulation. We compare the effect of multimodal and text-only clarifying questions in both tasks within a conversational search context from various perspectives. Our findings reveal that while participants showed a strong preference for multimodal questions when answering clarifying questions, preferences were more balanced in the query reformulation task. The impact of images varied with both task type and user expertise. In answering clarifying questions, images helped maintain engagement across different expertise levels, while in query reformulation they led to more precise queries and improved retrieval performance. Interestingly, for clarifying question answering, text-only setups demonstrated better user performance as they provided more comprehensive textual information in the absence of images. These results provide valuable insights for designing effective multimodal conversational search systems, highlighting that the benefits of visual augmentation are task-dependent and should be strategically implemented based on the specific search context and user characteristics.
format Preprint
id arxiv_https___arxiv_org_abs_2602_08700
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Do Images Clarify? A Study on the Effect of Images on Clarifying Questions in Conversational Search
Siro, Clemencia
Abbasiantaeb, Zahra
Yuan, Yifei
Aliannejadi, Mohammad
de Rijke, Maarten
Computation and Language
Human-Computer Interaction
Information Retrieval
Conversational search systems increasingly employ clarifying questions to refine user queries and improve the search experience. Previous studies have demonstrated the usefulness of text-based clarifying questions in enhancing both retrieval performance and user experience. While images have been shown to improve retrieval performance in various contexts, their impact on user performance when incorporated into clarifying questions remains largely unexplored. We conduct a user study with 73 participants to investigate the role of images in conversational search, specifically examining their effects on two search-related tasks: (i) answering clarifying questions and (ii) query reformulation. We compare the effect of multimodal and text-only clarifying questions in both tasks within a conversational search context from various perspectives. Our findings reveal that while participants showed a strong preference for multimodal questions when answering clarifying questions, preferences were more balanced in the query reformulation task. The impact of images varied with both task type and user expertise. In answering clarifying questions, images helped maintain engagement across different expertise levels, while in query reformulation they led to more precise queries and improved retrieval performance. Interestingly, for clarifying question answering, text-only setups demonstrated better user performance as they provided more comprehensive textual information in the absence of images. These results provide valuable insights for designing effective multimodal conversational search systems, highlighting that the benefits of visual augmentation are task-dependent and should be strategically implemented based on the specific search context and user characteristics.
title Do Images Clarify? A Study on the Effect of Images on Clarifying Questions in Conversational Search
topic Computation and Language
Human-Computer Interaction
Information Retrieval
url https://arxiv.org/abs/2602.08700