Learn from your own latents and not from tokens: A sample-complexity theory
Fuente:
arXiv
Guardado en:
| Autores principales: | Korchinski, Daniel J., Favero, Alessandro, Wyart, Matthieu |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Bigger Isn't Always Memorizing: Early Stopping Overparameterized Diffusion Models
por: Favero, Alessandro, et al.
Publicado: (2025)
por: Favero, Alessandro, et al.
Publicado: (2025)
On the Emergence of Linear Analogies in Word Embeddings
por: Korchinski, Daniel J., et al.
Publicado: (2025)
por: Korchinski, Daniel J., et al.
Publicado: (2025)
A Phase Transition in Diffusion Models Reveals the Hierarchical Nature of Data
por: Sclocchi, Antonio, et al.
Publicado: (2024)
por: Sclocchi, Antonio, et al.
Publicado: (2024)
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures
por: Cagnetta, Francesco, et al.
Publicado: (2025)
por: Cagnetta, Francesco, et al.
Publicado: (2025)
Symmetry in language statistics shapes the geometry of model representations
por: Karkada, Dhruva, et al.
Publicado: (2026)
por: Karkada, Dhruva, et al.
Publicado: (2026)
How Compositional Generalization and Creativity Improve as Diffusion Models are Trained
por: Favero, Alessandro, et al.
Publicado: (2025)
por: Favero, Alessandro, et al.
Publicado: (2025)
Sampling Data with Chains of Forward-Backward Diffusion Steps
por: Kang, Hyunmo, et al.
Publicado: (2026)
por: Kang, Hyunmo, et al.
Publicado: (2026)
Probing the Latent Hierarchical Structure of Data via Diffusion Models
por: Sclocchi, Antonio, et al.
Publicado: (2024)
por: Sclocchi, Antonio, et al.
Publicado: (2024)
How Deep Neural Networks Learn Compositional Data: The Random Hierarchy Model
por: Cagnetta, Francesco, et al.
Publicado: (2023)
por: Cagnetta, Francesco, et al.
Publicado: (2023)
Towards a theory of how the structure of language is acquired by deep neural networks
por: Cagnetta, Francesco, et al.
Publicado: (2024)
por: Cagnetta, Francesco, et al.
Publicado: (2024)
Learning curves theory for hierarchically compositional data with power-law distributed features
por: Cagnetta, Francesco, et al.
Publicado: (2025)
por: Cagnetta, Francesco, et al.
Publicado: (2025)
Hierarchical Concept Geometry in Language Models Emerges from Word Co-occurrence
por: Nava, Andres, et al.
Publicado: (2026)
por: Nava, Andres, et al.
Publicado: (2026)
How Deep Networks Learn Sparse and Hierarchical Data: the Sparse Random Hierarchy Model
por: Tomasini, Umberto, et al.
Publicado: (2024)
por: Tomasini, Umberto, et al.
Publicado: (2024)
On the different regimes of Stochastic Gradient Descent
por: Sclocchi, Antonio, et al.
Publicado: (2023)
por: Sclocchi, Antonio, et al.
Publicado: (2023)
Microscopic description of the intermittent dynamics driving logarithmic creep
por: Korchinski, Daniel J., et al.
Publicado: (2024)
por: Korchinski, Daniel J., et al.
Publicado: (2024)
Deep networks learn to parse uniform-depth context-free languages from local statistics
por: Parley, Jack T., et al.
Publicado: (2026)
por: Parley, Jack T., et al.
Publicado: (2026)
Deriving Neural Scaling Laws from the statistics of natural language
por: Cagnetta, Francesco, et al.
Publicado: (2026)
por: Cagnetta, Francesco, et al.
Publicado: (2026)
The Physics of Data and Tasks: Theories of Locality and Compositionality in Deep Learning
por: Favero, Alessandro
Publicado: (2025)
por: Favero, Alessandro
Publicado: (2025)
Diffusion Models Preferentially Memorize Prototypical Examples or: Why Does My Diffusion Model Love Slop?
por: Rodriguez, Marta Aparicio, et al.
Publicado: (2026)
por: Rodriguez, Marta Aparicio, et al.
Publicado: (2026)
Unified Latents (UL): How to train your latents
por: Heek, Jonathan, et al.
Publicado: (2026)
por: Heek, Jonathan, et al.
Publicado: (2026)
Task Addition and Weight Disentanglement in Closed-Vocabulary Models
por: Hazimeh, Adam, et al.
Publicado: (2025)
por: Hazimeh, Adam, et al.
Publicado: (2025)
Not all tokens are needed(NAT): token efficient reinforcement learning
por: Sang, Hejian, et al.
Publicado: (2026)
por: Sang, Hejian, et al.
Publicado: (2026)
Deep graph matching meets mixed-integer linear programming: Relax at your own risk ?
por: Xu, Zhoubo, et al.
Publicado: (2021)
por: Xu, Zhoubo, et al.
Publicado: (2021)
Efficient numeracy in language models through single-token number embeddings
por: Kreitner, Linus, et al.
Publicado: (2025)
por: Kreitner, Linus, et al.
Publicado: (2025)
Km-scale dynamical downscaling through conformalized latent diffusion models
por: Brusaferri, Alessandro, et al.
Publicado: (2025)
por: Brusaferri, Alessandro, et al.
Publicado: (2025)
Scaling FP8 training to trillion-token LLMs
por: Fishman, Maxim, et al.
Publicado: (2024)
por: Fishman, Maxim, et al.
Publicado: (2024)
Looking beyond the next token
por: Thankaraj, Abitha, et al.
Publicado: (2025)
por: Thankaraj, Abitha, et al.
Publicado: (2025)
Where is the signal in tokenization space?
por: Geh, Renato Lui, et al.
Publicado: (2024)
por: Geh, Renato Lui, et al.
Publicado: (2024)
Physics in Next-token Prediction
por: An, Hongjun, et al.
Publicado: (2024)
por: An, Hongjun, et al.
Publicado: (2024)
Unified token representations for sequential decision models
por: Tian, Zhuojing, et al.
Publicado: (2025)
por: Tian, Zhuojing, et al.
Publicado: (2025)
Graph2text or Graph2token: A Perspective of Large Language Models for Graph Learning
por: Yu, Shuo, et al.
Publicado: (2025)
por: Yu, Shuo, et al.
Publicado: (2025)
The pitfalls of next-token prediction
por: Bachmann, Gregor, et al.
Publicado: (2024)
por: Bachmann, Gregor, et al.
Publicado: (2024)
Non-asymptotic Convergence of Training Transformers for Next-token Prediction
por: Huang, Ruiquan, et al.
Publicado: (2024)
por: Huang, Ruiquan, et al.
Publicado: (2024)
GQ-VAE: A gated quantized VAE for learning variable length tokens
por: Datta, Theo, et al.
Publicado: (2025)
por: Datta, Theo, et al.
Publicado: (2025)
On the Stability of Iterative Retraining of Generative Models on their own Data
por: Bertrand, Quentin, et al.
Publicado: (2023)
por: Bertrand, Quentin, et al.
Publicado: (2023)
Next-token pretraining implies in-context learning
por: Riechers, Paul M., et al.
Publicado: (2025)
por: Riechers, Paul M., et al.
Publicado: (2025)
On multi-token prediction for efficient LLM inference
por: Mehra, Somesh, et al.
Publicado: (2025)
por: Mehra, Somesh, et al.
Publicado: (2025)
Fast, memory-efficient genomic interval tokenizers for modern machine learning
por: LeRoy, Nathan J., et al.
Publicado: (2025)
por: LeRoy, Nathan J., et al.
Publicado: (2025)
Hierarchical self-assembly for high-yield addressable complexity at fixed conditions
por: Holmes-Cerfon, Miranda, et al.
Publicado: (2025)
por: Holmes-Cerfon, Miranda, et al.
Publicado: (2025)
Tokenphormer: Structure-aware Multi-token Graph Transformer for Node Classification
por: Zhou, Zijie, et al.
Publicado: (2024)
por: Zhou, Zijie, et al.
Publicado: (2024)
Ejemplares similares
-
Bigger Isn't Always Memorizing: Early Stopping Overparameterized Diffusion Models
por: Favero, Alessandro, et al.
Publicado: (2025) -
On the Emergence of Linear Analogies in Word Embeddings
por: Korchinski, Daniel J., et al.
Publicado: (2025) -
A Phase Transition in Diffusion Models Reveals the Hierarchical Nature of Data
por: Sclocchi, Antonio, et al.
Publicado: (2024) -
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures
por: Cagnetta, Francesco, et al.
Publicado: (2025) -
Symmetry in language statistics shapes the geometry of model representations
por: Karkada, Dhruva, et al.
Publicado: (2026)