Analyzing the Presentation, Content, and Utilization of References in LLM-powered Conversational AI Systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ouyang, Jianheng, Narechania, Arpit
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914480396435456
author Ouyang, Jianheng
Narechania, Arpit
author_facet Ouyang, Jianheng
Narechania, Arpit
contents As conversational AI systems become popular for information retrieval and question-answering, the references they cite are key to ensuring their answers are reliable and trustworthy. Yet, no prior work systematically analyzes how these references are presented or their quality. We examine 1,517 references from 30 question-answer pairs across nine systems, focusing on their (1) presentation in the user interface and (2) quality using the CRAAP criteria. We find notable variations in the presentation, quality, and quantity of references across systems. For instance, ChatGPT provides more references (9.5 per response on average) with higher quality (15.48/20 CRAAP score), while Hunyuan-TurboS provides fewer references (4.0) and lower quality (11.65/20). Additionally, a preliminary user study shows that people rarely interact with these references and that their behavior differs across systems. These findings highlight the need for better interface designs that help users engage with and trust references more effectively.
format Preprint
id arxiv_https___arxiv_org_abs_2604_15326
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Analyzing the Presentation, Content, and Utilization of References in LLM-powered Conversational AI Systems
Ouyang, Jianheng
Narechania, Arpit
Human-Computer Interaction
As conversational AI systems become popular for information retrieval and question-answering, the references they cite are key to ensuring their answers are reliable and trustworthy. Yet, no prior work systematically analyzes how these references are presented or their quality. We examine 1,517 references from 30 question-answer pairs across nine systems, focusing on their (1) presentation in the user interface and (2) quality using the CRAAP criteria. We find notable variations in the presentation, quality, and quantity of references across systems. For instance, ChatGPT provides more references (9.5 per response on average) with higher quality (15.48/20 CRAAP score), while Hunyuan-TurboS provides fewer references (4.0) and lower quality (11.65/20). Additionally, a preliminary user study shows that people rarely interact with these references and that their behavior differs across systems. These findings highlight the need for better interface designs that help users engage with and trust references more effectively.
title Analyzing the Presentation, Content, and Utilization of References in LLM-powered Conversational AI Systems
topic Human-Computer Interaction
url https://arxiv.org/abs/2604.15326