On the Emergence of Linear Analogies in Word Embeddings
Fuente:
arXiv
Saved in:
| Main Authors: | Korchinski, Daniel J., Karkada, Dhruva, Bahri, Yasaman, Wyart, Matthieu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Symmetry in language statistics shapes the geometry of model representations
by: Karkada, Dhruva, et al.
Published: (2026)
by: Karkada, Dhruva, et al.
Published: (2026)
Towards a theory of how the structure of language is acquired by deep neural networks
by: Cagnetta, Francesco, et al.
Published: (2024)
by: Cagnetta, Francesco, et al.
Published: (2024)
Sampling Data with Chains of Forward-Backward Diffusion Steps
by: Kang, Hyunmo, et al.
Published: (2026)
by: Kang, Hyunmo, et al.
Published: (2026)
Deep networks learn to parse uniform-depth context-free languages from local statistics
by: Parley, Jack T., et al.
Published: (2026)
by: Parley, Jack T., et al.
Published: (2026)
On the different regimes of Stochastic Gradient Descent
by: Sclocchi, Antonio, et al.
Published: (2023)
by: Sclocchi, Antonio, et al.
Published: (2023)
How Deep Networks Learn Sparse and Hierarchical Data: the Sparse Random Hierarchy Model
by: Tomasini, Umberto, et al.
Published: (2024)
by: Tomasini, Umberto, et al.
Published: (2024)
Learning curves theory for hierarchically compositional data with power-law distributed features
by: Cagnetta, Francesco, et al.
Published: (2025)
by: Cagnetta, Francesco, et al.
Published: (2025)
Microscopic description of the intermittent dynamics driving logarithmic creep
by: Korchinski, Daniel J., et al.
Published: (2024)
by: Korchinski, Daniel J., et al.
Published: (2024)
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures
by: Cagnetta, Francesco, et al.
Published: (2025)
by: Cagnetta, Francesco, et al.
Published: (2025)
Probing the Latent Hierarchical Structure of Data via Diffusion Models
by: Sclocchi, Antonio, et al.
Published: (2024)
by: Sclocchi, Antonio, et al.
Published: (2024)
Explaining Neural Scaling Laws
by: Bahri, Yasaman, et al.
Published: (2021)
by: Bahri, Yasaman, et al.
Published: (2021)
A Phase Transition in Diffusion Models Reveals the Hierarchical Nature of Data
by: Sclocchi, Antonio, et al.
Published: (2024)
by: Sclocchi, Antonio, et al.
Published: (2024)
Analog Physical Systems Can Exhibit Double Descent
by: Dillavou, Sam, et al.
Published: (2025)
by: Dillavou, Sam, et al.
Published: (2025)
Emergence of Distortions in High-Dimensional Guided Diffusion Models
by: Ventura, Enrico, et al.
Published: (2026)
by: Ventura, Enrico, et al.
Published: (2026)
Exact Learning Dynamics of In-Context Learning in Linear Transformers and Its Application to Non-Linear Transformers
by: Mainali, Nischal, et al.
Published: (2025)
by: Mainali, Nischal, et al.
Published: (2025)
Short-range depinning in the presence of velocity-weakening
by: de Geus, Tom W. J., et al.
Published: (2024)
by: de Geus, Tom W. J., et al.
Published: (2024)
Dynamical heterogeneities of thermal creep in pinned interfaces
by: de Geus, Tom W. J., et al.
Published: (2024)
by: de Geus, Tom W. J., et al.
Published: (2024)
Learning Linear Regression with Low-Rank Tasks in-Context
by: Takanami, Kaito, et al.
Published: (2025)
by: Takanami, Kaito, et al.
Published: (2025)
Two-Point Deterministic Equivalence for Stochastic Gradient Dynamics in Linear Models
by: Atanasov, Alexander, et al.
Published: (2025)
by: Atanasov, Alexander, et al.
Published: (2025)
Closed-Form Training Dynamics Reveal Learned Features and Linear Structure in Word2Vec-like Models
by: Karkada, Dhruva, et al.
Published: (2025)
by: Karkada, Dhruva, et al.
Published: (2025)
Pruning-induced phases in fully-connected neural networks: the eumentia, the dementia, and the amentia
by: Pan, Haining, et al.
Published: (2026)
by: Pan, Haining, et al.
Published: (2026)
Precise Dynamics of Diagonal Linear Networks: A Unifying Analysis by Dynamical Mean-Field Theory
by: Nishiyama, Sota, et al.
Published: (2025)
by: Nishiyama, Sota, et al.
Published: (2025)
Properties of the geometry of solutions and capacity of multi-layer neural networks with Rectified Linear Units activations
by: Baldassi, Carlo, et al.
Published: (2019)
by: Baldassi, Carlo, et al.
Published: (2019)
Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer
by: Bordelon, Blake, et al.
Published: (2025)
by: Bordelon, Blake, et al.
Published: (2025)
Mapping of attention mechanisms to a generalized Potts model
by: Rende, Riccardo, et al.
Published: (2023)
by: Rende, Riccardo, et al.
Published: (2023)
How transformers learn structured data: insights from hierarchical filtering
by: Garnier-Brun, Jerome, et al.
Published: (2024)
by: Garnier-Brun, Jerome, et al.
Published: (2024)
Recurrent Self-Attention Dynamics: An Energy-Agnostic Perspective from Jacobians
by: Tomihari, Akiyoshi, et al.
Published: (2025)
by: Tomihari, Akiyoshi, et al.
Published: (2025)
Dropout Universality: Scaling Laws and Optimal Scheduling at the Edge-of-Chaos
by: Sarmiento, Lucas Fernandez
Published: (2026)
by: Sarmiento, Lucas Fernandez
Published: (2026)
Random Features Hopfield Networks generalize retrieval to previously unseen examples
by: Kalaj, Silvio, et al.
Published: (2024)
by: Kalaj, Silvio, et al.
Published: (2024)
Thermal Robustness of Retrieval in Dense Associative Memories: LSE vs LSR Kernels
by: Petrova, Tatiana
Published: (2026)
by: Petrova, Tatiana
Published: (2026)
Supervised Hebbian Learning
by: Alemanno, Francesco, et al.
Published: (2022)
by: Alemanno, Francesco, et al.
Published: (2022)
Prototype Analysis in Hopfield Networks with Hebbian Learning
by: McAlister, Hayden, et al.
Published: (2024)
by: McAlister, Hayden, et al.
Published: (2024)
A Generative Diffusion Model for Amorphous Materials
by: Yang, Kai, et al.
Published: (2025)
by: Yang, Kai, et al.
Published: (2025)
Soft Quantization: Model Compression Via Weight Coupling
by: Bernstein, Daniel T., et al.
Published: (2026)
by: Bernstein, Daniel T., et al.
Published: (2026)
How noise affects memory in linear recurrent networks
by: Guan, JingChuan, et al.
Published: (2024)
by: Guan, JingChuan, et al.
Published: (2024)
Generative modeling through internal high-dimensional chaotic activity
by: Fournier, Samantha J., et al.
Published: (2024)
by: Fournier, Samantha J., et al.
Published: (2024)
The Quantization Model of Neural Scaling
by: Michaud, Eric J., et al.
Published: (2023)
by: Michaud, Eric J., et al.
Published: (2023)
Grokking as the Transition from Lazy to Rich Training Dynamics
by: Kumar, Tanishq, et al.
Published: (2023)
by: Kumar, Tanishq, et al.
Published: (2023)
Finite-time Lyapunov exponents of deep neural networks
by: Storm, L., et al.
Published: (2023)
by: Storm, L., et al.
Published: (2023)
Exploring the Energy Landscape of RBMs: Reciprocal Space Insights into Bosons, Hierarchical Learning and Symmetry Breaking
by: Toledo-Marin, J. Quetzalcóatl, et al.
Published: (2025)
by: Toledo-Marin, J. Quetzalcóatl, et al.
Published: (2025)
Similar Items
-
Symmetry in language statistics shapes the geometry of model representations
by: Karkada, Dhruva, et al.
Published: (2026) -
Towards a theory of how the structure of language is acquired by deep neural networks
by: Cagnetta, Francesco, et al.
Published: (2024) -
Sampling Data with Chains of Forward-Backward Diffusion Steps
by: Kang, Hyunmo, et al.
Published: (2026) -
Deep networks learn to parse uniform-depth context-free languages from local statistics
by: Parley, Jack T., et al.
Published: (2026) -
On the different regimes of Stochastic Gradient Descent
by: Sclocchi, Antonio, et al.
Published: (2023)