Highlight All the Phrases: Enhancing LLM Transparency through Visual Factuality Indicators

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Do, Hyo Jin, Ostrand, Rachel, Geyer, Werner, Murugesan, Keerthiram, Wei, Dennis, Weisz, Justin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915437285998592
author Do, Hyo Jin
Ostrand, Rachel
Geyer, Werner
Murugesan, Keerthiram
Wei, Dennis
Weisz, Justin
author_facet Do, Hyo Jin
Ostrand, Rachel
Geyer, Werner
Murugesan, Keerthiram
Wei, Dennis
Weisz, Justin
contents Large language models (LLMs) are susceptible to generating inaccurate or false information, often referred to as "hallucinations" or "confabulations." While several technical advancements have been made to detect hallucinated content by assessing the factuality of the model's responses, there is still limited research on how to effectively communicate this information to users. To address this gap, we conducted two scenario-based experiments with a total of 208 participants to systematically compare the effects of various design strategies for communicating factuality scores by assessing participants' ratings of trust, ease in validating response accuracy, and preference. Our findings reveal that participants preferred and trusted a design in which all phrases within a response were color-coded based on factuality scores. Participants also found it easier to validate accuracy of the response in this style compared to a baseline with no style applied. Our study offers practical design guidelines for LLM application developers and designers, aimed at calibrating user trust, aligning with user preferences, and enhancing users' ability to scrutinize LLM outputs.
format Preprint
id arxiv_https___arxiv_org_abs_2508_06846
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Highlight All the Phrases: Enhancing LLM Transparency through Visual Factuality Indicators
Do, Hyo Jin
Ostrand, Rachel
Geyer, Werner
Murugesan, Keerthiram
Wei, Dennis
Weisz, Justin
Human-Computer Interaction
Artificial Intelligence
Large language models (LLMs) are susceptible to generating inaccurate or false information, often referred to as "hallucinations" or "confabulations." While several technical advancements have been made to detect hallucinated content by assessing the factuality of the model's responses, there is still limited research on how to effectively communicate this information to users. To address this gap, we conducted two scenario-based experiments with a total of 208 participants to systematically compare the effects of various design strategies for communicating factuality scores by assessing participants' ratings of trust, ease in validating response accuracy, and preference. Our findings reveal that participants preferred and trusted a design in which all phrases within a response were color-coded based on factuality scores. Participants also found it easier to validate accuracy of the response in this style compared to a baseline with no style applied. Our study offers practical design guidelines for LLM application developers and designers, aimed at calibrating user trust, aligning with user preferences, and enhancing users' ability to scrutinize LLM outputs.
title Highlight All the Phrases: Enhancing LLM Transparency through Visual Factuality Indicators
topic Human-Computer Interaction
Artificial Intelligence
url https://arxiv.org/abs/2508.06846