Correlation Dimension of Natural Language in a Statistical Manifold
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866917666433794048 |
|---|---|
| author | Du, Xin Tanaka-Ishii, Kumiko |
| author_facet | Du, Xin Tanaka-Ishii, Kumiko |
| contents | The correlation dimension of natural language is measured by applying the Grassberger-Procaccia algorithm to high-dimensional sequences produced by a large-scale language model. This method, previously studied only in a Euclidean space, is reformulated in a statistical manifold via the Fisher-Rao distance. Language exhibits a multifractal, with global self-similarity and a universal dimension around 6.5, which is smaller than those of simple discrete random sequences and larger than that of a Barabási-Albert process. Long memory is the key to producing self-similarity. Our method is applicable to any probabilistic model of real-world discrete sequences, and we show an application to music data. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2405_06321 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Correlation Dimension of Natural Language in a Statistical Manifold Du, Xin Tanaka-Ishii, Kumiko Computation and Language Statistical Mechanics Artificial Intelligence The correlation dimension of natural language is measured by applying the Grassberger-Procaccia algorithm to high-dimensional sequences produced by a large-scale language model. This method, previously studied only in a Euclidean space, is reformulated in a statistical manifold via the Fisher-Rao distance. Language exhibits a multifractal, with global self-similarity and a universal dimension around 6.5, which is smaller than those of simple discrete random sequences and larger than that of a Barabási-Albert process. Long memory is the key to producing self-similarity. Our method is applicable to any probabilistic model of real-world discrete sequences, and we show an application to music data. |
| title | Correlation Dimension of Natural Language in a Statistical Manifold |
| topic | Computation and Language Statistical Mechanics Artificial Intelligence |
| url | https://arxiv.org/abs/2405.06321 |