ATLANTIS at SemEval-2025 Task 3: Detecting Hallucinated Text Spans in Question Answering

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kobus, Catherine, Lancelot, François, Martin, Marion-Cécile, Amer, Nawal Ould
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908480968851456
author Kobus, Catherine
Lancelot, François
Martin, Marion-Cécile
Amer, Nawal Ould
author_facet Kobus, Catherine
Lancelot, François
Martin, Marion-Cécile
Amer, Nawal Ould
contents This paper presents the contributions of the ATLANTIS team to SemEval-2025 Task 3, focusing on detecting hallucinated text spans in question answering systems. Large Language Models (LLMs) have significantly advanced Natural Language Generation (NLG) but remain susceptible to hallucinations, generating incorrect or misleading content. To address this, we explored methods both with and without external context, utilizing few-shot prompting with a LLM, token-level classification or LLM fine-tuned on synthetic data. Notably, our approaches achieved top rankings in Spanish and competitive placements in English and German. This work highlights the importance of integrating relevant context to mitigate hallucinations and demonstrate the potential of fine-tuned models and prompt engineering.
format Preprint
id arxiv_https___arxiv_org_abs_2508_05179
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ATLANTIS at SemEval-2025 Task 3: Detecting Hallucinated Text Spans in Question Answering
Kobus, Catherine
Lancelot, François
Martin, Marion-Cécile
Amer, Nawal Ould
Computation and Language
This paper presents the contributions of the ATLANTIS team to SemEval-2025 Task 3, focusing on detecting hallucinated text spans in question answering systems. Large Language Models (LLMs) have significantly advanced Natural Language Generation (NLG) but remain susceptible to hallucinations, generating incorrect or misleading content. To address this, we explored methods both with and without external context, utilizing few-shot prompting with a LLM, token-level classification or LLM fine-tuned on synthetic data. Notably, our approaches achieved top rankings in Spanish and competitive placements in English and German. This work highlights the importance of integrating relevant context to mitigate hallucinations and demonstrate the potential of fine-tuned models and prompt engineering.
title ATLANTIS at SemEval-2025 Task 3: Detecting Hallucinated Text Spans in Question Answering
topic Computation and Language
url https://arxiv.org/abs/2508.05179