Assessing Web Search Credibility and Response Groundedness in Chat Assistants

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Vykopal, Ivan, Pikuliak, Matúš, Ostermann, Simon, Šimko, Marián
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908840225669120
author Vykopal, Ivan
Pikuliak, Matúš
Ostermann, Simon
Šimko, Marián
author_facet Vykopal, Ivan
Pikuliak, Matúš
Ostermann, Simon
Šimko, Marián
contents Chat assistants increasingly integrate web search functionality, enabling them to retrieve and cite external sources. While this promises more reliable answers, it also raises the risk of amplifying misinformation from low-credibility sources. In this paper, we introduce a novel methodology for evaluating assistants' web search behavior, focusing on source credibility and the groundedness of responses with respect to cited sources. Using 100 claims across five misinformation-prone topics, we assess GPT-4o, GPT-5, Perplexity, and Qwen Chat. Our findings reveal differences between the assistants, with Perplexity achieving the highest source credibility, whereas GPT-4o exhibits elevated citation of non-credibility sources on sensitive topics. This work provides the first systematic comparison of commonly used chat assistants for fact-checking behavior, offering a foundation for evaluating AI systems in high-stakes information environments.
format Preprint
id arxiv_https___arxiv_org_abs_2510_13749
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Assessing Web Search Credibility and Response Groundedness in Chat Assistants
Vykopal, Ivan
Pikuliak, Matúš
Ostermann, Simon
Šimko, Marián
Computation and Language
Chat assistants increasingly integrate web search functionality, enabling them to retrieve and cite external sources. While this promises more reliable answers, it also raises the risk of amplifying misinformation from low-credibility sources. In this paper, we introduce a novel methodology for evaluating assistants' web search behavior, focusing on source credibility and the groundedness of responses with respect to cited sources. Using 100 claims across five misinformation-prone topics, we assess GPT-4o, GPT-5, Perplexity, and Qwen Chat. Our findings reveal differences between the assistants, with Perplexity achieving the highest source credibility, whereas GPT-4o exhibits elevated citation of non-credibility sources on sensitive topics. This work provides the first systematic comparison of commonly used chat assistants for fact-checking behavior, offering a foundation for evaluating AI systems in high-stakes information environments.
title Assessing Web Search Credibility and Response Groundedness in Chat Assistants
topic Computation and Language
url https://arxiv.org/abs/2510.13749