Testing Transformer Learnability on the Arithmetic Sequence of Rooted Trees
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915647196233728 |
|---|---|
| author | Breccia, Alessandro Gerace, Federica Lippi, Marco Sicuro, Gabriele Contucci, Pierluigi |
| author_facet | Breccia, Alessandro Gerace, Federica Lippi, Marco Sicuro, Gabriele Contucci, Pierluigi |
| contents | We study whether a Large Language Model can learn the deterministic sequence of trees generated by the iterated prime factorization of the natural numbers. Each integer is mapped into a rooted planar tree and the resulting sequence $ \mathbb{N}\mathcal{T}$ defines an arithmetic text with measurable statistical structure. A transformer network (the GPT-2 architecture) is trained from scratch on the first $10^{11}$ elements to subsequently test its predictive ability under next-word and masked-word prediction tasks. Our results show that the model partially learns the internal grammar of $\mathbb{N}\mathcal{T}$, capturing non-trivial regularities and correlations. This suggests that learnability may extend beyond empirical data to the very structure of arithmetic. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_01870 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Testing Transformer Learnability on the Arithmetic Sequence of Rooted Trees Breccia, Alessandro Gerace, Federica Lippi, Marco Sicuro, Gabriele Contucci, Pierluigi Artificial Intelligence Disordered Systems and Neural Networks Mathematical Physics Number Theory We study whether a Large Language Model can learn the deterministic sequence of trees generated by the iterated prime factorization of the natural numbers. Each integer is mapped into a rooted planar tree and the resulting sequence $ \mathbb{N}\mathcal{T}$ defines an arithmetic text with measurable statistical structure. A transformer network (the GPT-2 architecture) is trained from scratch on the first $10^{11}$ elements to subsequently test its predictive ability under next-word and masked-word prediction tasks. Our results show that the model partially learns the internal grammar of $\mathbb{N}\mathcal{T}$, capturing non-trivial regularities and correlations. This suggests that learnability may extend beyond empirical data to the very structure of arithmetic. |
| title | Testing Transformer Learnability on the Arithmetic Sequence of Rooted Trees |
| topic | Artificial Intelligence Disordered Systems and Neural Networks Mathematical Physics Number Theory |
| url | https://arxiv.org/abs/2512.01870 |