Deep networks learn to parse uniform-depth context-free languages from local statistics
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Parley, Jack T., Cagnetta, Francesco, Wyart, Matthieu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards a theory of how the structure of language is acquired by deep neural networks
von: Cagnetta, Francesco, et al.
Veröffentlicht: (2024)
von: Cagnetta, Francesco, et al.
Veröffentlicht: (2024)
Learning curves theory for hierarchically compositional data with power-law distributed features
von: Cagnetta, Francesco, et al.
Veröffentlicht: (2025)
von: Cagnetta, Francesco, et al.
Veröffentlicht: (2025)
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures
von: Cagnetta, Francesco, et al.
Veröffentlicht: (2025)
von: Cagnetta, Francesco, et al.
Veröffentlicht: (2025)
Symmetry in language statistics shapes the geometry of model representations
von: Karkada, Dhruva, et al.
Veröffentlicht: (2026)
von: Karkada, Dhruva, et al.
Veröffentlicht: (2026)
How Deep Networks Learn Sparse and Hierarchical Data: the Sparse Random Hierarchy Model
von: Tomasini, Umberto, et al.
Veröffentlicht: (2024)
von: Tomasini, Umberto, et al.
Veröffentlicht: (2024)
On the Emergence of Linear Analogies in Word Embeddings
von: Korchinski, Daniel J., et al.
Veröffentlicht: (2025)
von: Korchinski, Daniel J., et al.
Veröffentlicht: (2025)
On the different regimes of Stochastic Gradient Descent
von: Sclocchi, Antonio, et al.
Veröffentlicht: (2023)
von: Sclocchi, Antonio, et al.
Veröffentlicht: (2023)
Probing the Latent Hierarchical Structure of Data via Diffusion Models
von: Sclocchi, Antonio, et al.
Veröffentlicht: (2024)
von: Sclocchi, Antonio, et al.
Veröffentlicht: (2024)
A Phase Transition in Diffusion Models Reveals the Hierarchical Nature of Data
von: Sclocchi, Antonio, et al.
Veröffentlicht: (2024)
von: Sclocchi, Antonio, et al.
Veröffentlicht: (2024)
Sampling Data with Chains of Forward-Backward Diffusion Steps
von: Kang, Hyunmo, et al.
Veröffentlicht: (2026)
von: Kang, Hyunmo, et al.
Veröffentlicht: (2026)
Applying statistical learning theory to deep learning
von: Gerbelot, Cédric, et al.
Veröffentlicht: (2023)
von: Gerbelot, Cédric, et al.
Veröffentlicht: (2023)
Asymptotic theory of in-context learning by linear attention
von: Lu, Yue M., et al.
Veröffentlicht: (2024)
von: Lu, Yue M., et al.
Veröffentlicht: (2024)
Implicit bias produces neural scaling laws in learning curves, from perceptrons to deep networks
von: D'Amico, Francesco, et al.
Veröffentlicht: (2025)
von: D'Amico, Francesco, et al.
Veröffentlicht: (2025)
High-dimensional learning of narrow neural networks
von: Cui, Hugo
Veröffentlicht: (2024)
von: Cui, Hugo
Veröffentlicht: (2024)
Self-attention as an attractor network: transient memories without backpropagation
von: D'Amico, Francesco, et al.
Veröffentlicht: (2024)
von: D'Amico, Francesco, et al.
Veröffentlicht: (2024)
The committee machine: Computational to statistical gaps in learning a two-layers neural network
von: Aubin, Benjamin, et al.
Veröffentlicht: (2018)
von: Aubin, Benjamin, et al.
Veröffentlicht: (2018)
Deep neural networks from the perspective of ergodic theory
von: Zhang, Fan
Veröffentlicht: (2023)
von: Zhang, Fan
Veröffentlicht: (2023)
Injectivity of ReLU networks: perspectives from statistical physics
von: Maillard, Antoine, et al.
Veröffentlicht: (2023)
von: Maillard, Antoine, et al.
Veröffentlicht: (2023)
How transformers learn structured data: insights from hierarchical filtering
von: Garnier-Brun, Jerome, et al.
Veröffentlicht: (2024)
von: Garnier-Brun, Jerome, et al.
Veröffentlicht: (2024)
Asymptotics of feature learning in two-layer networks after one gradient-step
von: Cui, Hugo, et al.
Veröffentlicht: (2024)
von: Cui, Hugo, et al.
Veröffentlicht: (2024)
Adaptive kernel predictors from feature-learning infinite limits of neural networks
von: Lauditi, Clarissa, et al.
Veröffentlicht: (2025)
von: Lauditi, Clarissa, et al.
Veröffentlicht: (2025)
Ductile and brittle yielding of athermal amorphous solids: a mean-field paradigm beyond the random field Ising model
von: Parley, Jack T., et al.
Veröffentlicht: (2023)
von: Parley, Jack T., et al.
Veröffentlicht: (2023)
Pruning-induced phases in fully-connected neural networks: the eumentia, the dementia, and the amentia
von: Pan, Haining, et al.
Veröffentlicht: (2026)
von: Pan, Haining, et al.
Veröffentlicht: (2026)
A statistical physics framework for optimal learning
von: Mignacco, Francesca, et al.
Veröffentlicht: (2025)
von: Mignacco, Francesca, et al.
Veröffentlicht: (2025)
Sharp description of local minima in the loss landscape of high-dimensional two-layer ReLU neural networks
von: Huang, Jie, et al.
Veröffentlicht: (2026)
von: Huang, Jie, et al.
Veröffentlicht: (2026)
Distinct mechanisms underlying in-context learning in transformers
von: Gibson, Cole, et al.
Veröffentlicht: (2026)
von: Gibson, Cole, et al.
Veröffentlicht: (2026)
Fundamental limits of learning in sequence multi-index models and deep attention networks: High-dimensional asymptotics and sharp thresholds
von: Troiani, Emanuele, et al.
Veröffentlicht: (2025)
von: Troiani, Emanuele, et al.
Veröffentlicht: (2025)
Supervised Hebbian Learning
von: Alemanno, Francesco, et al.
Veröffentlicht: (2022)
von: Alemanno, Francesco, et al.
Veröffentlicht: (2022)
Neural Langevin Machine: a local asymmetric learning rule can be creative
von: Yu, Zhendong, et al.
Veröffentlicht: (2025)
von: Yu, Zhendong, et al.
Veröffentlicht: (2025)
Generative diffusion for perceptron problems: statistical physics analysis and efficient algorithms
von: Demyanenko, Elizaveta, et al.
Veröffentlicht: (2025)
von: Demyanenko, Elizaveta, et al.
Veröffentlicht: (2025)
How noise affects memory in linear recurrent networks
von: Guan, JingChuan, et al.
Veröffentlicht: (2024)
von: Guan, JingChuan, et al.
Veröffentlicht: (2024)
Differential learning kinetics govern the transition from memorization to generalization during in-context learning
von: Nguyen, Alex, et al.
Veröffentlicht: (2024)
von: Nguyen, Alex, et al.
Veröffentlicht: (2024)
In-context denoising with one-layer transformers: connections between attention and associative memory retrieval
von: Smart, Matthew, et al.
Veröffentlicht: (2025)
von: Smart, Matthew, et al.
Veröffentlicht: (2025)
Exact full-RSB SAT/UNSAT transition in infinitely wide two-layer neural networks
von: Annesi, Brandon L., et al.
Veröffentlicht: (2024)
von: Annesi, Brandon L., et al.
Veröffentlicht: (2024)
Generalization emerges from local optimization in a self-organized learning network
von: Barland, S., et al.
Veröffentlicht: (2024)
von: Barland, S., et al.
Veröffentlicht: (2024)
Theory of Optimal Learning Rate Schedules and Scaling Laws for a Random Feature Model
von: Bordelon, Blake, et al.
Veröffentlicht: (2026)
von: Bordelon, Blake, et al.
Veröffentlicht: (2026)
Dimension-free deterministic equivalents and scaling laws for random feature regression
von: Defilippis, Leonardo, et al.
Veröffentlicht: (2024)
von: Defilippis, Leonardo, et al.
Veröffentlicht: (2024)
Emergent weight morphologies in deep neural networks
von: de Jong, Pascal, et al.
Veröffentlicht: (2025)
von: de Jong, Pascal, et al.
Veröffentlicht: (2025)
A universal approximation theorem for nonlinear resistive networks
von: Scellier, Benjamin, et al.
Veröffentlicht: (2023)
von: Scellier, Benjamin, et al.
Veröffentlicht: (2023)
Supervised and Unsupervised protocols for hetero-associative neural networks
von: Alessandrelli, Andrea, et al.
Veröffentlicht: (2025)
von: Alessandrelli, Andrea, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Towards a theory of how the structure of language is acquired by deep neural networks
von: Cagnetta, Francesco, et al.
Veröffentlicht: (2024) -
Learning curves theory for hierarchically compositional data with power-law distributed features
von: Cagnetta, Francesco, et al.
Veröffentlicht: (2025) -
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures
von: Cagnetta, Francesco, et al.
Veröffentlicht: (2025) -
Symmetry in language statistics shapes the geometry of model representations
von: Karkada, Dhruva, et al.
Veröffentlicht: (2026) -
How Deep Networks Learn Sparse and Hierarchical Data: the Sparse Random Hierarchy Model
von: Tomasini, Umberto, et al.
Veröffentlicht: (2024)