Will It Still Be True Tomorrow? Multilingual Evergreen Question Classification to Improve Trustworthy QA

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Pletenev, Sergey, Marina, Maria, Ivanov, Nikolay, Galimzianova, Daria, Krayko, Nikita, Salnikov, Mikhail, Konovalov, Vasily, Panchenko, Alexander, Moskvoretskii, Viktor
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912397901430784
author Pletenev, Sergey
Marina, Maria
Ivanov, Nikolay
Galimzianova, Daria
Krayko, Nikita
Salnikov, Mikhail
Konovalov, Vasily
Panchenko, Alexander
Moskvoretskii, Viktor
author_facet Pletenev, Sergey
Marina, Maria
Ivanov, Nikolay
Galimzianova, Daria
Krayko, Nikita
Salnikov, Mikhail
Konovalov, Vasily
Panchenko, Alexander
Moskvoretskii, Viktor
contents Large Language Models (LLMs) often hallucinate in question answering (QA) tasks. A key yet underexplored factor contributing to this is the temporality of questions -- whether they are evergreen (answers remain stable over time) or mutable (answers change). In this work, we introduce EverGreenQA, the first multilingual QA dataset with evergreen labels, supporting both evaluation and training. Using EverGreenQA, we benchmark 12 modern LLMs to assess whether they encode question temporality explicitly (via verbalized judgments) or implicitly (via uncertainty signals). We also train EG-E5, a lightweight multilingual classifier that achieves SoTA performance on this task. Finally, we demonstrate the practical utility of evergreen classification across three applications: improving self-knowledge estimation, filtering QA datasets, and explaining GPT-4o retrieval behavior.
format Preprint
id arxiv_https___arxiv_org_abs_2505_21115
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Will It Still Be True Tomorrow? Multilingual Evergreen Question Classification to Improve Trustworthy QA
Pletenev, Sergey
Marina, Maria
Ivanov, Nikolay
Galimzianova, Daria
Krayko, Nikita
Salnikov, Mikhail
Konovalov, Vasily
Panchenko, Alexander
Moskvoretskii, Viktor
Computation and Language
Large Language Models (LLMs) often hallucinate in question answering (QA) tasks. A key yet underexplored factor contributing to this is the temporality of questions -- whether they are evergreen (answers remain stable over time) or mutable (answers change). In this work, we introduce EverGreenQA, the first multilingual QA dataset with evergreen labels, supporting both evaluation and training. Using EverGreenQA, we benchmark 12 modern LLMs to assess whether they encode question temporality explicitly (via verbalized judgments) or implicitly (via uncertainty signals). We also train EG-E5, a lightweight multilingual classifier that achieves SoTA performance on this task. Finally, we demonstrate the practical utility of evergreen classification across three applications: improving self-knowledge estimation, filtering QA datasets, and explaining GPT-4o retrieval behavior.
title Will It Still Be True Tomorrow? Multilingual Evergreen Question Classification to Improve Trustworthy QA
topic Computation and Language
url https://arxiv.org/abs/2505.21115