Testing Transformer Learnability on the Arithmetic Sequence of Rooted Trees

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Breccia, Alessandro, Gerace, Federica, Lippi, Marco, Sicuro, Gabriele, Contucci, Pierluigi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915647196233728
author Breccia, Alessandro
Gerace, Federica
Lippi, Marco
Sicuro, Gabriele
Contucci, Pierluigi
author_facet Breccia, Alessandro
Gerace, Federica
Lippi, Marco
Sicuro, Gabriele
Contucci, Pierluigi
contents We study whether a Large Language Model can learn the deterministic sequence of trees generated by the iterated prime factorization of the natural numbers. Each integer is mapped into a rooted planar tree and the resulting sequence $ \mathbb{N}\mathcal{T}$ defines an arithmetic text with measurable statistical structure. A transformer network (the GPT-2 architecture) is trained from scratch on the first $10^{11}$ elements to subsequently test its predictive ability under next-word and masked-word prediction tasks. Our results show that the model partially learns the internal grammar of $\mathbb{N}\mathcal{T}$, capturing non-trivial regularities and correlations. This suggests that learnability may extend beyond empirical data to the very structure of arithmetic.
format Preprint
id arxiv_https___arxiv_org_abs_2512_01870
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Testing Transformer Learnability on the Arithmetic Sequence of Rooted Trees
Breccia, Alessandro
Gerace, Federica
Lippi, Marco
Sicuro, Gabriele
Contucci, Pierluigi
Artificial Intelligence
Disordered Systems and Neural Networks
Mathematical Physics
Number Theory
We study whether a Large Language Model can learn the deterministic sequence of trees generated by the iterated prime factorization of the natural numbers. Each integer is mapped into a rooted planar tree and the resulting sequence $ \mathbb{N}\mathcal{T}$ defines an arithmetic text with measurable statistical structure. A transformer network (the GPT-2 architecture) is trained from scratch on the first $10^{11}$ elements to subsequently test its predictive ability under next-word and masked-word prediction tasks. Our results show that the model partially learns the internal grammar of $\mathbb{N}\mathcal{T}$, capturing non-trivial regularities and correlations. This suggests that learnability may extend beyond empirical data to the very structure of arithmetic.
title Testing Transformer Learnability on the Arithmetic Sequence of Rooted Trees
topic Artificial Intelligence
Disordered Systems and Neural Networks
Mathematical Physics
Number Theory
url https://arxiv.org/abs/2512.01870