Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Fernández-Pichel, Marcos, Pichel, Juan C., Losada, David E.
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2407.12468
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909526171582464
author Fernández-Pichel, Marcos
Pichel, Juan C.
Losada, David E.
author_facet Fernández-Pichel, Marcos
Pichel, Juan C.
Losada, David E.
contents Search engines (SEs) have traditionally been primary tools for information seeking, but the new Large Language Models (LLMs) are emerging as powerful alternatives, particularly for question-answering tasks. This study compares the performance of four popular SEs, seven LLMs, and retrieval-augmented (RAG) variants in answering 150 health-related questions from the TREC Health Misinformation (HM) Track. Results reveal SEs correctly answer between 50 and 70% of questions, often hindered by many retrieval results not responding to the health question. LLMs deliver higher accuracy, correctly answering about 80% of questions, though their performance is sensitive to input prompts. RAG methods significantly enhance smaller LLMs' effectiveness, improving accuracy by up to 30% by integrating retrieval evidence.
format Preprint
id arxiv_https___arxiv_org_abs_2407_12468
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Evaluating Search Engines and Large Language Models for Answering Health Questions
Fernández-Pichel, Marcos
Pichel, Juan C.
Losada, David E.
Information Retrieval
Artificial Intelligence
Search engines (SEs) have traditionally been primary tools for information seeking, but the new Large Language Models (LLMs) are emerging as powerful alternatives, particularly for question-answering tasks. This study compares the performance of four popular SEs, seven LLMs, and retrieval-augmented (RAG) variants in answering 150 health-related questions from the TREC Health Misinformation (HM) Track. Results reveal SEs correctly answer between 50 and 70% of questions, often hindered by many retrieval results not responding to the health question. LLMs deliver higher accuracy, correctly answering about 80% of questions, though their performance is sensitive to input prompts. RAG methods significantly enhance smaller LLMs' effectiveness, improving accuracy by up to 30% by integrating retrieval evidence.
title Evaluating Search Engines and Large Language Models for Answering Health Questions
topic Information Retrieval
Artificial Intelligence
url https://arxiv.org/abs/2407.12468