When Facts Change: Probing LLMs on Evolving Knowledge with evolveQA

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Nakshatri, Nishanth Sridhar, Roy, Shamik, Arivazhagan, Manoj Ghuhan, Zhou, Hanhan, Kumar, Vinayshekhar Bannihatti, Gangadharaiah, Rashmi
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909905066131456
author Nakshatri, Nishanth Sridhar
Roy, Shamik
Arivazhagan, Manoj Ghuhan
Zhou, Hanhan
Kumar, Vinayshekhar Bannihatti
Gangadharaiah, Rashmi
author_facet Nakshatri, Nishanth Sridhar
Roy, Shamik
Arivazhagan, Manoj Ghuhan
Zhou, Hanhan
Kumar, Vinayshekhar Bannihatti
Gangadharaiah, Rashmi
contents LLMs often fail to handle temporal knowledge conflicts--contradictions arising when facts evolve over time within their training data. Existing studies evaluate this phenomenon through benchmarks built on structured knowledge bases like Wikidata, but they focus on widely-covered, easily-memorized popular entities and lack the dynamic structure needed to fairly evaluate LLMs with different knowledge cut-off dates. We introduce evolveQA, a benchmark specifically designed to evaluate LLMs on temporally evolving knowledge, constructed from 3 real-world, time-stamped corpora: AWS updates, Azure changes, and WHO disease outbreak reports. Our framework identifies naturally occurring knowledge evolution and generates questions with gold answers tailored to different LLM knowledge cut-off dates. Through extensive evaluation of 12 open and closed-source LLMs across 3 knowledge probing formats, we demonstrate significant performance drops of up to 31% on evolveQA compared to static knowledge questions.
format Preprint
id arxiv_https___arxiv_org_abs_2510_19172
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle When Facts Change: Probing LLMs on Evolving Knowledge with evolveQA
Nakshatri, Nishanth Sridhar
Roy, Shamik
Arivazhagan, Manoj Ghuhan
Zhou, Hanhan
Kumar, Vinayshekhar Bannihatti
Gangadharaiah, Rashmi
Computation and Language
Artificial Intelligence
LLMs often fail to handle temporal knowledge conflicts--contradictions arising when facts evolve over time within their training data. Existing studies evaluate this phenomenon through benchmarks built on structured knowledge bases like Wikidata, but they focus on widely-covered, easily-memorized popular entities and lack the dynamic structure needed to fairly evaluate LLMs with different knowledge cut-off dates. We introduce evolveQA, a benchmark specifically designed to evaluate LLMs on temporally evolving knowledge, constructed from 3 real-world, time-stamped corpora: AWS updates, Azure changes, and WHO disease outbreak reports. Our framework identifies naturally occurring knowledge evolution and generates questions with gold answers tailored to different LLM knowledge cut-off dates. Through extensive evaluation of 12 open and closed-source LLMs across 3 knowledge probing formats, we demonstrate significant performance drops of up to 31% on evolveQA compared to static knowledge questions.
title When Facts Change: Probing LLMs on Evolving Knowledge with evolveQA
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.19172