ACL-Verbatim: hallucination-free question answering for research

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Recski, Gábor, Tóth, Szilveszter, Verdha, Nadia, Boros, István, Kovács, Ádám
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914583121231872
author Recski, Gábor
Tóth, Szilveszter
Verdha, Nadia
Boros, István
Kovács, Ádám
author_facet Recski, Gábor
Tóth, Szilveszter
Verdha, Nadia
Boros, István
Kovács, Ádám
contents Academic researchers need efficient and reliable methods for collecting high-quality information from trusted sources, but modern tools for AI-assisted research still suffer from the tendency of Large Language Models (LLMs) to produce factually inaccurate or nonsensical output, commonly referred to as hallucinations. We apply the extractive question answering system VerbatimRAG to research papers in the ACL Anthology, directly mapping user queries to verbatim text spans in retrieved documents. We contribute a novel ground truth dataset for the task of mapping user queries to relevant text spans in research papers, and use it to train and evaluate a variety of extractive models. Human annotation is performed by NLP researchers and is based on synthetic user queries generated using a custom pipeline based on the ScIRGen methodology, paired with chunks of research papers retrieved by VerbatimRAG. On this benchmark, a 150M-parameter ModernBERT token classifier trained on silver supervision from our pipeline achieves the best word-level F1 (53.6), ahead of the strongest evaluated LLM extractor (48.7).
format Preprint
id arxiv_https___arxiv_org_abs_2605_21102
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ACL-Verbatim: hallucination-free question answering for research
Recski, Gábor
Tóth, Szilveszter
Verdha, Nadia
Boros, István
Kovács, Ádám
Computation and Language
Artificial Intelligence
Software Engineering
Academic researchers need efficient and reliable methods for collecting high-quality information from trusted sources, but modern tools for AI-assisted research still suffer from the tendency of Large Language Models (LLMs) to produce factually inaccurate or nonsensical output, commonly referred to as hallucinations. We apply the extractive question answering system VerbatimRAG to research papers in the ACL Anthology, directly mapping user queries to verbatim text spans in retrieved documents. We contribute a novel ground truth dataset for the task of mapping user queries to relevant text spans in research papers, and use it to train and evaluate a variety of extractive models. Human annotation is performed by NLP researchers and is based on synthetic user queries generated using a custom pipeline based on the ScIRGen methodology, paired with chunks of research papers retrieved by VerbatimRAG. On this benchmark, a 150M-parameter ModernBERT token classifier trained on silver supervision from our pipeline achieves the best word-level F1 (53.6), ahead of the strongest evaluated LLM extractor (48.7).
title ACL-Verbatim: hallucination-free question answering for research
topic Computation and Language
Artificial Intelligence
Software Engineering
url https://arxiv.org/abs/2605.21102