Will It Still Be True Tomorrow? Multilingual Evergreen Question Classification to Improve Trustworthy QA
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866912397901430784 |
|---|---|
| author | Pletenev, Sergey Marina, Maria Ivanov, Nikolay Galimzianova, Daria Krayko, Nikita Salnikov, Mikhail Konovalov, Vasily Panchenko, Alexander Moskvoretskii, Viktor |
| author_facet | Pletenev, Sergey Marina, Maria Ivanov, Nikolay Galimzianova, Daria Krayko, Nikita Salnikov, Mikhail Konovalov, Vasily Panchenko, Alexander Moskvoretskii, Viktor |
| contents | Large Language Models (LLMs) often hallucinate in question answering (QA) tasks. A key yet underexplored factor contributing to this is the temporality of questions -- whether they are evergreen (answers remain stable over time) or mutable (answers change). In this work, we introduce EverGreenQA, the first multilingual QA dataset with evergreen labels, supporting both evaluation and training. Using EverGreenQA, we benchmark 12 modern LLMs to assess whether they encode question temporality explicitly (via verbalized judgments) or implicitly (via uncertainty signals). We also train EG-E5, a lightweight multilingual classifier that achieves SoTA performance on this task. Finally, we demonstrate the practical utility of evergreen classification across three applications: improving self-knowledge estimation, filtering QA datasets, and explaining GPT-4o retrieval behavior. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_21115 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Will It Still Be True Tomorrow? Multilingual Evergreen Question Classification to Improve Trustworthy QA Pletenev, Sergey Marina, Maria Ivanov, Nikolay Galimzianova, Daria Krayko, Nikita Salnikov, Mikhail Konovalov, Vasily Panchenko, Alexander Moskvoretskii, Viktor Computation and Language Large Language Models (LLMs) often hallucinate in question answering (QA) tasks. A key yet underexplored factor contributing to this is the temporality of questions -- whether they are evergreen (answers remain stable over time) or mutable (answers change). In this work, we introduce EverGreenQA, the first multilingual QA dataset with evergreen labels, supporting both evaluation and training. Using EverGreenQA, we benchmark 12 modern LLMs to assess whether they encode question temporality explicitly (via verbalized judgments) or implicitly (via uncertainty signals). We also train EG-E5, a lightweight multilingual classifier that achieves SoTA performance on this task. Finally, we demonstrate the practical utility of evergreen classification across three applications: improving self-knowledge estimation, filtering QA datasets, and explaining GPT-4o retrieval behavior. |
| title | Will It Still Be True Tomorrow? Multilingual Evergreen Question Classification to Improve Trustworthy QA |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2505.21115 |