Guardado en:
Detalles Bibliográficos
Autores principales: Tschisgale, Paul, Wulff, Peter
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:https://arxiv.org/abs/2602.15889
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910111653429248
author Tschisgale, Paul
Wulff, Peter
author_facet Tschisgale, Paul
Wulff, Peter
contents Large language models (LLMs) are increasingly used in research as both tools and objects of study. Much of this work assumes that LLM performance under fixed conditions (identical model snapshot, hyperparameters, and prompt) is time-invariant, meaning that average output quality remains stable over time; otherwise, reliability and reproducibility would be compromised. To test the assumption of time invariance, we conducted a longitudinal study of GPT-4o's average performance under fixed conditions. The LLM was queried to solve the same physics task ten times every three hours over approximately three months. Spectral (Fourier) analysis of the resulting time series revealed substantial periodic variability, accounting for about 20% of total variance. The observed periodic patterns are consistent with interacting daily and weekly rhythms. These findings challenge the assumption of time invariance and carry important implications for research involving LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2602_15889
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Daily and Weekly Periodicity in Large Language Model Performance and Its Implications for Research
Tschisgale, Paul
Wulff, Peter
Applications
Artificial Intelligence
Computation and Language
Physics Education
Large language models (LLMs) are increasingly used in research as both tools and objects of study. Much of this work assumes that LLM performance under fixed conditions (identical model snapshot, hyperparameters, and prompt) is time-invariant, meaning that average output quality remains stable over time; otherwise, reliability and reproducibility would be compromised. To test the assumption of time invariance, we conducted a longitudinal study of GPT-4o's average performance under fixed conditions. The LLM was queried to solve the same physics task ten times every three hours over approximately three months. Spectral (Fourier) analysis of the resulting time series revealed substantial periodic variability, accounting for about 20% of total variance. The observed periodic patterns are consistent with interacting daily and weekly rhythms. These findings challenge the assumption of time invariance and carry important implications for research involving LLMs.
title Daily and Weekly Periodicity in Large Language Model Performance and Its Implications for Research
topic Applications
Artificial Intelligence
Computation and Language
Physics Education
url https://arxiv.org/abs/2602.15889