Saved in:
| Main Authors: | Korchinski, Daniel J., Favero, Alessandro, Wyart, Matthieu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2605.27734 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bigger Isn't Always Memorizing: Early Stopping Overparameterized Diffusion Models
by: Favero, Alessandro, et al.
Published: (2025)
by: Favero, Alessandro, et al.
Published: (2025)
On the Emergence of Linear Analogies in Word Embeddings
by: Korchinski, Daniel J., et al.
Published: (2025)
by: Korchinski, Daniel J., et al.
Published: (2025)
A Phase Transition in Diffusion Models Reveals the Hierarchical Nature of Data
by: Sclocchi, Antonio, et al.
Published: (2024)
by: Sclocchi, Antonio, et al.
Published: (2024)
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures
by: Cagnetta, Francesco, et al.
Published: (2025)
by: Cagnetta, Francesco, et al.
Published: (2025)
Symmetry in language statistics shapes the geometry of model representations
by: Karkada, Dhruva, et al.
Published: (2026)
by: Karkada, Dhruva, et al.
Published: (2026)
Sampling Data with Chains of Forward-Backward Diffusion Steps
by: Kang, Hyunmo, et al.
Published: (2026)
by: Kang, Hyunmo, et al.
Published: (2026)
How Compositional Generalization and Creativity Improve as Diffusion Models are Trained
by: Favero, Alessandro, et al.
Published: (2025)
by: Favero, Alessandro, et al.
Published: (2025)
Probing the Latent Hierarchical Structure of Data via Diffusion Models
by: Sclocchi, Antonio, et al.
Published: (2024)
by: Sclocchi, Antonio, et al.
Published: (2024)
How Deep Neural Networks Learn Compositional Data: The Random Hierarchy Model
by: Cagnetta, Francesco, et al.
Published: (2023)
by: Cagnetta, Francesco, et al.
Published: (2023)
Towards a theory of how the structure of language is acquired by deep neural networks
by: Cagnetta, Francesco, et al.
Published: (2024)
by: Cagnetta, Francesco, et al.
Published: (2024)
Learning curves theory for hierarchically compositional data with power-law distributed features
by: Cagnetta, Francesco, et al.
Published: (2025)
by: Cagnetta, Francesco, et al.
Published: (2025)
Hierarchical Concept Geometry in Language Models Emerges from Word Co-occurrence
by: Nava, Andres, et al.
Published: (2026)
by: Nava, Andres, et al.
Published: (2026)
How Deep Networks Learn Sparse and Hierarchical Data: the Sparse Random Hierarchy Model
by: Tomasini, Umberto, et al.
Published: (2024)
by: Tomasini, Umberto, et al.
Published: (2024)
On the different regimes of Stochastic Gradient Descent
by: Sclocchi, Antonio, et al.
Published: (2023)
by: Sclocchi, Antonio, et al.
Published: (2023)
Microscopic description of the intermittent dynamics driving logarithmic creep
by: Korchinski, Daniel J., et al.
Published: (2024)
by: Korchinski, Daniel J., et al.
Published: (2024)
Deep networks learn to parse uniform-depth context-free languages from local statistics
by: Parley, Jack T., et al.
Published: (2026)
by: Parley, Jack T., et al.
Published: (2026)
Deriving Neural Scaling Laws from the statistics of natural language
by: Cagnetta, Francesco, et al.
Published: (2026)
by: Cagnetta, Francesco, et al.
Published: (2026)
The Physics of Data and Tasks: Theories of Locality and Compositionality in Deep Learning
by: Favero, Alessandro
Published: (2025)
by: Favero, Alessandro
Published: (2025)
Diffusion Models Preferentially Memorize Prototypical Examples or: Why Does My Diffusion Model Love Slop?
by: Rodriguez, Marta Aparicio, et al.
Published: (2026)
by: Rodriguez, Marta Aparicio, et al.
Published: (2026)
Unified Latents (UL): How to train your latents
by: Heek, Jonathan, et al.
Published: (2026)
by: Heek, Jonathan, et al.
Published: (2026)
Task Addition and Weight Disentanglement in Closed-Vocabulary Models
by: Hazimeh, Adam, et al.
Published: (2025)
by: Hazimeh, Adam, et al.
Published: (2025)
Deep graph matching meets mixed-integer linear programming: Relax at your own risk ?
by: Xu, Zhoubo, et al.
Published: (2021)
by: Xu, Zhoubo, et al.
Published: (2021)
Not all tokens are needed(NAT): token efficient reinforcement learning
by: Sang, Hejian, et al.
Published: (2026)
by: Sang, Hejian, et al.
Published: (2026)
Efficient numeracy in language models through single-token number embeddings
by: Kreitner, Linus, et al.
Published: (2025)
by: Kreitner, Linus, et al.
Published: (2025)
Km-scale dynamical downscaling through conformalized latent diffusion models
by: Brusaferri, Alessandro, et al.
Published: (2025)
by: Brusaferri, Alessandro, et al.
Published: (2025)
Scaling FP8 training to trillion-token LLMs
by: Fishman, Maxim, et al.
Published: (2024)
by: Fishman, Maxim, et al.
Published: (2024)
Looking beyond the next token
by: Thankaraj, Abitha, et al.
Published: (2025)
by: Thankaraj, Abitha, et al.
Published: (2025)
Where is the signal in tokenization space?
by: Geh, Renato Lui, et al.
Published: (2024)
by: Geh, Renato Lui, et al.
Published: (2024)
Physics in Next-token Prediction
by: An, Hongjun, et al.
Published: (2024)
by: An, Hongjun, et al.
Published: (2024)
Hierarchical self-assembly for high-yield addressable complexity at fixed conditions
by: Holmes-Cerfon, Miranda, et al.
Published: (2025)
by: Holmes-Cerfon, Miranda, et al.
Published: (2025)
Unified token representations for sequential decision models
by: Tian, Zhuojing, et al.
Published: (2025)
by: Tian, Zhuojing, et al.
Published: (2025)
Backdoor Unlearning by Linear Task Decomposition
by: Abdelraheem, Amel, et al.
Published: (2025)
by: Abdelraheem, Amel, et al.
Published: (2025)
The pitfalls of next-token prediction
by: Bachmann, Gregor, et al.
Published: (2024)
by: Bachmann, Gregor, et al.
Published: (2024)
MEMOIR: Lifelong Model Editing with Minimal Overwrite and Informed Retention for LLMs
by: Wang, Ke, et al.
Published: (2025)
by: Wang, Ke, et al.
Published: (2025)
Graph2text or Graph2token: A Perspective of Large Language Models for Graph Learning
by: Yu, Shuo, et al.
Published: (2025)
by: Yu, Shuo, et al.
Published: (2025)
Did the Neurons Read your Book? Document-level Membership Inference for Large Language Models
by: Meeus, Matthieu, et al.
Published: (2023)
by: Meeus, Matthieu, et al.
Published: (2023)
Next-token pretraining implies in-context learning
by: Riechers, Paul M., et al.
Published: (2025)
by: Riechers, Paul M., et al.
Published: (2025)
On multi-token prediction for efficient LLM inference
by: Mehra, Somesh, et al.
Published: (2025)
by: Mehra, Somesh, et al.
Published: (2025)
On the Stability of Iterative Retraining of Generative Models on their own Data
by: Bertrand, Quentin, et al.
Published: (2023)
by: Bertrand, Quentin, et al.
Published: (2023)
Non-asymptotic Convergence of Training Transformers for Next-token Prediction
by: Huang, Ruiquan, et al.
Published: (2024)
by: Huang, Ruiquan, et al.
Published: (2024)
Similar Items
-
Bigger Isn't Always Memorizing: Early Stopping Overparameterized Diffusion Models
by: Favero, Alessandro, et al.
Published: (2025) -
On the Emergence of Linear Analogies in Word Embeddings
by: Korchinski, Daniel J., et al.
Published: (2025) -
A Phase Transition in Diffusion Models Reveals the Hierarchical Nature of Data
by: Sclocchi, Antonio, et al.
Published: (2024) -
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures
by: Cagnetta, Francesco, et al.
Published: (2025) -
Symmetry in language statistics shapes the geometry of model representations
by: Karkada, Dhruva, et al.
Published: (2026)