Search Engines in an AI Era: The False Promise of Factual and Verifiable Source-Cited Responses

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Venkit, Pranav Narayanan, Laban, Philippe, Zhou, Yilun, Mao, Yixin, Wu, Chien-Sheng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909370828193792
author Venkit, Pranav Narayanan
Laban, Philippe
Zhou, Yilun
Mao, Yixin
Wu, Chien-Sheng
author_facet Venkit, Pranav Narayanan
Laban, Philippe
Zhou, Yilun
Mao, Yixin
Wu, Chien-Sheng
contents Large Language Model (LLM)-based applications are graduating from research prototypes to products serving millions of users, influencing how people write and consume information. A prominent example is the appearance of Answer Engines: LLM-based generative search engines supplanting traditional search engines. Answer engines not only retrieve relevant sources to a user query but synthesize answer summaries that cite the sources. To understand these systems' limitations, we first conducted a study with 21 participants, evaluating interactions with answer vs. traditional search engines and identifying 16 answer engine limitations. From these insights, we propose 16 answer engine design recommendations, linked to 8 metrics. An automated evaluation implementing our metrics on three popular engines (You.com, Perplexity.ai, BingChat) quantifies common limitations (e.g., frequent hallucination, inaccurate citation) and unique features (e.g., variation in answer confidence), with results mirroring user study insights. We release our Answer Engine Evaluation benchmark (AEE) to facilitate transparent evaluation of LLM-based applications.
format Preprint
id arxiv_https___arxiv_org_abs_2410_22349
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Search Engines in an AI Era: The False Promise of Factual and Verifiable Source-Cited Responses
Venkit, Pranav Narayanan
Laban, Philippe
Zhou, Yilun
Mao, Yixin
Wu, Chien-Sheng
Information Retrieval
Artificial Intelligence
Computation and Language
Computers and Society
Human-Computer Interaction
Large Language Model (LLM)-based applications are graduating from research prototypes to products serving millions of users, influencing how people write and consume information. A prominent example is the appearance of Answer Engines: LLM-based generative search engines supplanting traditional search engines. Answer engines not only retrieve relevant sources to a user query but synthesize answer summaries that cite the sources. To understand these systems' limitations, we first conducted a study with 21 participants, evaluating interactions with answer vs. traditional search engines and identifying 16 answer engine limitations. From these insights, we propose 16 answer engine design recommendations, linked to 8 metrics. An automated evaluation implementing our metrics on three popular engines (You.com, Perplexity.ai, BingChat) quantifies common limitations (e.g., frequent hallucination, inaccurate citation) and unique features (e.g., variation in answer confidence), with results mirroring user study insights. We release our Answer Engine Evaluation benchmark (AEE) to facilitate transparent evaluation of LLM-based applications.
title Search Engines in an AI Era: The False Promise of Factual and Verifiable Source-Cited Responses
topic Information Retrieval
Artificial Intelligence
Computation and Language
Computers and Society
Human-Computer Interaction
url https://arxiv.org/abs/2410.22349