Saved in:
Bibliographic Details
Main Authors: Chartier, Mathieu, Dakkoune, Nabil, Bourgeois, Guillaume, Jean, Stéphane
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2406.15173
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911928789499904
author Chartier, Mathieu
Dakkoune, Nabil
Bourgeois, Guillaume
Jean, Stéphane
author_facet Chartier, Mathieu
Dakkoune, Nabil
Bourgeois, Guillaume
Jean, Stéphane
contents Large Language Models (LLMs) like ChatGPT or Bard have revolutionized information retrieval and captivated the audience with their ability to generate custom responses in record time, regardless of the topic. In this article, we assess the capabilities of various LLMs in producing reliable, comprehensive, and sufficiently relevant responses about historical facts in French. To achieve this, we constructed a testbed comprising numerous history-related questions of varying types, themes, and levels of difficulty. Our evaluation of responses from ten selected LLMs reveals numerous shortcomings in both substance and form. Beyond an overall insufficient accuracy rate, we highlight uneven treatment of the French language, as well as issues related to verbosity and inconsistency in the responses provided by LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2406_15173
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Évaluation des capacités de réponse de larges modèles de langage (LLM) pour des questions d'historiens
Chartier, Mathieu
Dakkoune, Nabil
Bourgeois, Guillaume
Jean, Stéphane
Information Retrieval
Artificial Intelligence
Large Language Models (LLMs) like ChatGPT or Bard have revolutionized information retrieval and captivated the audience with their ability to generate custom responses in record time, regardless of the topic. In this article, we assess the capabilities of various LLMs in producing reliable, comprehensive, and sufficiently relevant responses about historical facts in French. To achieve this, we constructed a testbed comprising numerous history-related questions of varying types, themes, and levels of difficulty. Our evaluation of responses from ten selected LLMs reveals numerous shortcomings in both substance and form. Beyond an overall insufficient accuracy rate, we highlight uneven treatment of the French language, as well as issues related to verbosity and inconsistency in the responses provided by LLMs.
title Évaluation des capacités de réponse de larges modèles de langage (LLM) pour des questions d'historiens
topic Information Retrieval
Artificial Intelligence
url https://arxiv.org/abs/2406.15173