Saved in:
Bibliographic Details
Main Authors: Ou, Yulin, Wang, Yu, Xu, Yang, Buschmeier, Hendrik
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2510.27241
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911620519690240
author Ou, Yulin
Wang, Yu
Xu, Yang
Buschmeier, Hendrik
author_facet Ou, Yulin
Wang, Yu
Xu, Yang
Buschmeier, Hendrik
contents Recent theoretical advancement of information density in natural language has brought the following question on desk: To what degree does natural language exhibit periodicity pattern in its encoded information? We address this question by introducing a new method called AutoPeriod of Surprisal (APS). APS adopts a canonical periodicity detection algorithm and is able to identify any significant periods that exist in the surprisal sequence of a single document. By applying the algorithm to a set of corpora, we have obtained the following interesting results: Firstly, a considerable proportion of human language demonstrates a strong pattern of periodicity in information; Secondly, new periods that are outside the distributions of typical structural units in text (e.g., sentence boundaries, elementary discourse units, etc.) are found and further confirmed via harmonic regression modeling. We conclude that the periodicity of information in language is a joint outcome from both structured factors and other driving factors that take effect at longer distances. The advantages of our periodicity detection method and its potentials in LLM-generation detection are further discussed.
format Preprint
id arxiv_https___arxiv_org_abs_2510_27241
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Identifying the Periodicity of Information in Natural Language
Ou, Yulin
Wang, Yu
Xu, Yang
Buschmeier, Hendrik
Computation and Language
Recent theoretical advancement of information density in natural language has brought the following question on desk: To what degree does natural language exhibit periodicity pattern in its encoded information? We address this question by introducing a new method called AutoPeriod of Surprisal (APS). APS adopts a canonical periodicity detection algorithm and is able to identify any significant periods that exist in the surprisal sequence of a single document. By applying the algorithm to a set of corpora, we have obtained the following interesting results: Firstly, a considerable proportion of human language demonstrates a strong pattern of periodicity in information; Secondly, new periods that are outside the distributions of typical structural units in text (e.g., sentence boundaries, elementary discourse units, etc.) are found and further confirmed via harmonic regression modeling. We conclude that the periodicity of information in language is a joint outcome from both structured factors and other driving factors that take effect at longer distances. The advantages of our periodicity detection method and its potentials in LLM-generation detection are further discussed.
title Identifying the Periodicity of Information in Natural Language
topic Computation and Language
url https://arxiv.org/abs/2510.27241