Symmetry in language statistics shapes the geometry of model representations
Fuente:
arXiv
Salvato in:
| Autori principali: | Karkada, Dhruva, Korchinski, Daniel J., Nava, Andres, Wyart, Matthieu, Bahri, Yasaman |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
On the Emergence of Linear Analogies in Word Embeddings
di: Korchinski, Daniel J., et al.
Pubblicazione: (2025)
di: Korchinski, Daniel J., et al.
Pubblicazione: (2025)
Deep networks learn to parse uniform-depth context-free languages from local statistics
di: Parley, Jack T., et al.
Pubblicazione: (2026)
di: Parley, Jack T., et al.
Pubblicazione: (2026)
Towards a theory of how the structure of language is acquired by deep neural networks
di: Cagnetta, Francesco, et al.
Pubblicazione: (2024)
di: Cagnetta, Francesco, et al.
Pubblicazione: (2024)
Sampling Data with Chains of Forward-Backward Diffusion Steps
di: Kang, Hyunmo, et al.
Pubblicazione: (2026)
di: Kang, Hyunmo, et al.
Pubblicazione: (2026)
On the different regimes of Stochastic Gradient Descent
di: Sclocchi, Antonio, et al.
Pubblicazione: (2023)
di: Sclocchi, Antonio, et al.
Pubblicazione: (2023)
How Deep Networks Learn Sparse and Hierarchical Data: the Sparse Random Hierarchy Model
di: Tomasini, Umberto, et al.
Pubblicazione: (2024)
di: Tomasini, Umberto, et al.
Pubblicazione: (2024)
Learning curves theory for hierarchically compositional data with power-law distributed features
di: Cagnetta, Francesco, et al.
Pubblicazione: (2025)
di: Cagnetta, Francesco, et al.
Pubblicazione: (2025)
Microscopic description of the intermittent dynamics driving logarithmic creep
di: Korchinski, Daniel J., et al.
Pubblicazione: (2024)
di: Korchinski, Daniel J., et al.
Pubblicazione: (2024)
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures
di: Cagnetta, Francesco, et al.
Pubblicazione: (2025)
di: Cagnetta, Francesco, et al.
Pubblicazione: (2025)
Probing the Latent Hierarchical Structure of Data via Diffusion Models
di: Sclocchi, Antonio, et al.
Pubblicazione: (2024)
di: Sclocchi, Antonio, et al.
Pubblicazione: (2024)
How does training shape the Riemannian geometry of neural network representations?
di: Zavatone-Veth, Jacob A., et al.
Pubblicazione: (2023)
di: Zavatone-Veth, Jacob A., et al.
Pubblicazione: (2023)
Explaining Neural Scaling Laws
di: Bahri, Yasaman, et al.
Pubblicazione: (2021)
di: Bahri, Yasaman, et al.
Pubblicazione: (2021)
A Phase Transition in Diffusion Models Reveals the Hierarchical Nature of Data
di: Sclocchi, Antonio, et al.
Pubblicazione: (2024)
di: Sclocchi, Antonio, et al.
Pubblicazione: (2024)
Bias-inducing geometries: an exactly solvable data model with fairness implications
di: Mannelli, Stefano Sarao, et al.
Pubblicazione: (2022)
di: Mannelli, Stefano Sarao, et al.
Pubblicazione: (2022)
Applying statistical learning theory to deep learning
di: Gerbelot, Cédric, et al.
Pubblicazione: (2023)
di: Gerbelot, Cédric, et al.
Pubblicazione: (2023)
Exploring the Energy Landscape of RBMs: Reciprocal Space Insights into Bosons, Hierarchical Learning and Symmetry Breaking
di: Toledo-Marin, J. Quetzalcóatl, et al.
Pubblicazione: (2025)
di: Toledo-Marin, J. Quetzalcóatl, et al.
Pubblicazione: (2025)
Generative diffusion for perceptron problems: statistical physics analysis and efficient algorithms
di: Demyanenko, Elizaveta, et al.
Pubblicazione: (2025)
di: Demyanenko, Elizaveta, et al.
Pubblicazione: (2025)
Properties of the geometry of solutions and capacity of multi-layer neural networks with Rectified Linear Units activations
di: Baldassi, Carlo, et al.
Pubblicazione: (2019)
di: Baldassi, Carlo, et al.
Pubblicazione: (2019)
Short-range depinning in the presence of velocity-weakening
di: de Geus, Tom W. J., et al.
Pubblicazione: (2024)
di: de Geus, Tom W. J., et al.
Pubblicazione: (2024)
Sampling at intermediate temperatures is optimal for training large language models in protein structure prediction
di: Ghiringhelli, L., et al.
Pubblicazione: (2026)
di: Ghiringhelli, L., et al.
Pubblicazione: (2026)
Dynamical heterogeneities of thermal creep in pinned interfaces
di: de Geus, Tom W. J., et al.
Pubblicazione: (2024)
di: de Geus, Tom W. J., et al.
Pubblicazione: (2024)
Generalization through variance: how noise shapes inductive biases in diffusion models
di: Vastola, John J.
Pubblicazione: (2025)
di: Vastola, John J.
Pubblicazione: (2025)
Generative modeling through internal high-dimensional chaotic activity
di: Fournier, Samantha J., et al.
Pubblicazione: (2024)
di: Fournier, Samantha J., et al.
Pubblicazione: (2024)
Injectivity of ReLU networks: perspectives from statistical physics
di: Maillard, Antoine, et al.
Pubblicazione: (2023)
di: Maillard, Antoine, et al.
Pubblicazione: (2023)
Dynamical Regimes of Multimodal Diffusion Models
di: Albrychiewicz, Emil, et al.
Pubblicazione: (2026)
di: Albrychiewicz, Emil, et al.
Pubblicazione: (2026)
Autoregressive model path dependence near Ising criticality
di: Teoh, Yi Hong, et al.
Pubblicazione: (2024)
di: Teoh, Yi Hong, et al.
Pubblicazione: (2024)
Nadaraya-Watson kernel smoothing as a random energy model
di: Zavatone-Veth, Jacob A., et al.
Pubblicazione: (2024)
di: Zavatone-Veth, Jacob A., et al.
Pubblicazione: (2024)
A solvable model of learning generative diffusion: theory and insights
di: Cui, Hugo, et al.
Pubblicazione: (2025)
di: Cui, Hugo, et al.
Pubblicazione: (2025)
Non-equilibrium active noise enhances generative memory in diffusion models
di: Behera, Agnish Kumar, et al.
Pubblicazione: (2024)
di: Behera, Agnish Kumar, et al.
Pubblicazione: (2024)
A Federated Many-to-One Hopfield model for associative Neural Networks
di: Alessandrelli, Andrea, et al.
Pubblicazione: (2026)
di: Alessandrelli, Andrea, et al.
Pubblicazione: (2026)
A Generative Diffusion Model for Amorphous Materials
di: Yang, Kai, et al.
Pubblicazione: (2025)
di: Yang, Kai, et al.
Pubblicazione: (2025)
An exactly solvable model for emergence and scaling laws in the multitask sparse parity problem
di: Nam, Yoonsoo, et al.
Pubblicazione: (2024)
di: Nam, Yoonsoo, et al.
Pubblicazione: (2024)
Optimal thresholds and algorithms for a model of multi-modal learning in high dimensions
di: Keup, Christian, et al.
Pubblicazione: (2024)
di: Keup, Christian, et al.
Pubblicazione: (2024)
Specialization of softmax attention heads: insights from the high-dimensional single-location model
di: Sagitova, M., et al.
Pubblicazione: (2026)
di: Sagitova, M., et al.
Pubblicazione: (2026)
Statistical physics analysis of graph neural networks: Approaching optimality in the contextual stochastic block model
di: Duranthon, O., et al.
Pubblicazione: (2025)
di: Duranthon, O., et al.
Pubblicazione: (2025)
Soft Quantization: Model Compression Via Weight Coupling
di: Bernstein, Daniel T., et al.
Pubblicazione: (2026)
di: Bernstein, Daniel T., et al.
Pubblicazione: (2026)
Universal and nonuniversal statistics of transmission in thin random layered media
di: Park, Jongchul, et al.
Pubblicazione: (2022)
di: Park, Jongchul, et al.
Pubblicazione: (2022)
Mapping of attention mechanisms to a generalized Potts model
di: Rende, Riccardo, et al.
Pubblicazione: (2023)
di: Rende, Riccardo, et al.
Pubblicazione: (2023)
Fundamental limits of learning in sequence multi-index models and deep attention networks: High-dimensional asymptotics and sharp thresholds
di: Troiani, Emanuele, et al.
Pubblicazione: (2025)
di: Troiani, Emanuele, et al.
Pubblicazione: (2025)
A solvable high-dimensional model where nonlinear autoencoders learn structure invisible to PCA while test loss misaligns with generalization
di: Mendes, Vicente Conde, et al.
Pubblicazione: (2026)
di: Mendes, Vicente Conde, et al.
Pubblicazione: (2026)
Documenti analoghi
-
On the Emergence of Linear Analogies in Word Embeddings
di: Korchinski, Daniel J., et al.
Pubblicazione: (2025) -
Deep networks learn to parse uniform-depth context-free languages from local statistics
di: Parley, Jack T., et al.
Pubblicazione: (2026) -
Towards a theory of how the structure of language is acquired by deep neural networks
di: Cagnetta, Francesco, et al.
Pubblicazione: (2024) -
Sampling Data with Chains of Forward-Backward Diffusion Steps
di: Kang, Hyunmo, et al.
Pubblicazione: (2026) -
On the different regimes of Stochastic Gradient Descent
di: Sclocchi, Antonio, et al.
Pubblicazione: (2023)