Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2406.15173 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911928789499904 |
|---|---|
| author | Chartier, Mathieu Dakkoune, Nabil Bourgeois, Guillaume Jean, Stéphane |
| author_facet | Chartier, Mathieu Dakkoune, Nabil Bourgeois, Guillaume Jean, Stéphane |
| contents | Large Language Models (LLMs) like ChatGPT or Bard have revolutionized information retrieval and captivated the audience with their ability to generate custom responses in record time, regardless of the topic. In this article, we assess the capabilities of various LLMs in producing reliable, comprehensive, and sufficiently relevant responses about historical facts in French. To achieve this, we constructed a testbed comprising numerous history-related questions of varying types, themes, and levels of difficulty. Our evaluation of responses from ten selected LLMs reveals numerous shortcomings in both substance and form. Beyond an overall insufficient accuracy rate, we highlight uneven treatment of the French language, as well as issues related to verbosity and inconsistency in the responses provided by LLMs. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2406_15173 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Évaluation des capacités de réponse de larges modèles de langage (LLM) pour des questions d'historiens Chartier, Mathieu Dakkoune, Nabil Bourgeois, Guillaume Jean, Stéphane Information Retrieval Artificial Intelligence Large Language Models (LLMs) like ChatGPT or Bard have revolutionized information retrieval and captivated the audience with their ability to generate custom responses in record time, regardless of the topic. In this article, we assess the capabilities of various LLMs in producing reliable, comprehensive, and sufficiently relevant responses about historical facts in French. To achieve this, we constructed a testbed comprising numerous history-related questions of varying types, themes, and levels of difficulty. Our evaluation of responses from ten selected LLMs reveals numerous shortcomings in both substance and form. Beyond an overall insufficient accuracy rate, we highlight uneven treatment of the French language, as well as issues related to verbosity and inconsistency in the responses provided by LLMs. |
| title | Évaluation des capacités de réponse de larges modèles de langage (LLM) pour des questions d'historiens |
| topic | Information Retrieval Artificial Intelligence |
| url | https://arxiv.org/abs/2406.15173 |