Survival of the Fittest Representation: A Case Study with Modular Addition
Fuente:
arXiv
Saved in:
| Main Authors: | Ding, Xiaoman Delores, Guo, Zifan Carl, Michaud, Eric J., Liu, Ziming, Tegmark, Max |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Quantization Model of Neural Scaling
by: Michaud, Eric J., et al.
Published: (2023)
by: Michaud, Eric J., et al.
Published: (2023)
Physics of Skill Learning
by: Liu, Ziming, et al.
Published: (2025)
by: Liu, Ziming, et al.
Published: (2025)
On the creation of narrow AI: hierarchy and nonlocality of neural network skills
by: Michaud, Eric J., et al.
Published: (2025)
by: Michaud, Eric J., et al.
Published: (2025)
A Neural Scaling Law from Lottery Ticket Ensembling
by: Liu, Ziming, et al.
Published: (2023)
by: Liu, Ziming, et al.
Published: (2023)
Opening the AI black box: program synthesis via mechanistic interpretability
by: Michaud, Eric J., et al.
Published: (2024)
by: Michaud, Eric J., et al.
Published: (2024)
Investigating Representation Universality: Case Study on Genealogical Representations
by: Baek, David D., et al.
Published: (2024)
by: Baek, David D., et al.
Published: (2024)
OptPDE: Discovering Novel Integrable Systems via AI-Human Collaboration
by: Kantamneni, Subhash, et al.
Published: (2024)
by: Kantamneni, Subhash, et al.
Published: (2024)
Do Two AI Scientists Agree?
by: Fu, Xinghong, et al.
Published: (2025)
by: Fu, Xinghong, et al.
Published: (2025)
Language Models Use Trigonometry to Do Addition
by: Kantamneni, Subhash, et al.
Published: (2025)
by: Kantamneni, Subhash, et al.
Published: (2025)
GenEFT: Understanding Statics and Dynamics of Model Generalization via Effective Theory
by: Baek, David D., et al.
Published: (2024)
by: Baek, David D., et al.
Published: (2024)
Not All Language Model Features Are One-Dimensionally Linear
by: Engels, Joshua, et al.
Published: (2024)
by: Engels, Joshua, et al.
Published: (2024)
How Do Transformers "Do" Physics? Investigating the Simple Harmonic Oscillator
by: Kantamneni, Subhash, et al.
Published: (2024)
by: Kantamneni, Subhash, et al.
Published: (2024)
Towards Understanding Distilled Reasoning Models: A Representational Approach
by: Baek, David D., et al.
Published: (2025)
by: Baek, David D., et al.
Published: (2025)
Harmonic Loss Trains Interpretable AI Models
by: Baek, David D., et al.
Published: (2025)
by: Baek, David D., et al.
Published: (2025)
Efficient Dictionary Learning with Switch Sparse Autoencoders
by: Mudide, Anish, et al.
Published: (2024)
by: Mudide, Anish, et al.
Published: (2024)
A Resource Model For Neural Scaling Law
by: Song, Jinyeop, et al.
Published: (2024)
by: Song, Jinyeop, et al.
Published: (2024)
R.I.P.: Better Models by Survival of the Fittest Prompts
by: Yu, Ping, et al.
Published: (2025)
by: Yu, Ping, et al.
Published: (2025)
Neural Thermodynamic Laws for Large Language Model Training
by: Liu, Ziming, et al.
Published: (2025)
by: Liu, Ziming, et al.
Published: (2025)
The Odyssey of the Fittest: Can Agents Survive and Still Be Good?
by: Waldner, Dylan, et al.
Published: (2025)
by: Waldner, Dylan, et al.
Published: (2025)
Survival of the Fittest: Evolutionary Adaptation of Policies for Environmental Shifts
by: Paul, Sheryl, et al.
Published: (2024)
by: Paul, Sheryl, et al.
Published: (2024)
The Geometry of Concepts: Sparse Autoencoder Feature Structure
by: Li, Yuxiao, et al.
Published: (2024)
by: Li, Yuxiao, et al.
Published: (2024)
Are Sparse Autoencoders Useful? A Case Study in Sparse Probing
by: Kantamneni, Subhash, et al.
Published: (2025)
by: Kantamneni, Subhash, et al.
Published: (2025)
Grokking of Diffusion Models: Case Study on Modular Addition
by: Kim, Joon Hyeok, et al.
Published: (2026)
by: Kim, Joon Hyeok, et al.
Published: (2026)
KAN 2.0: Kolmogorov-Arnold Networks Meet Science
by: Liu, Ziming, et al.
Published: (2024)
by: Liu, Ziming, et al.
Published: (2024)
Language Models Represent Space and Time
by: Gurnee, Wes, et al.
Published: (2023)
by: Gurnee, Wes, et al.
Published: (2023)
Decomposing The Dark Matter of Sparse Autoencoders
by: Engels, Joshua, et al.
Published: (2024)
by: Engels, Joshua, et al.
Published: (2024)
Low-Rank Adapting Models for Sparse Autoencoders
by: Chen, Matthew, et al.
Published: (2025)
by: Chen, Matthew, et al.
Published: (2025)
Provable Benefits of Sinusoidal Activation for Modular Addition
by: Huang, Tianlong, et al.
Published: (2025)
by: Huang, Tianlong, et al.
Published: (2025)
Why Do You Grok? A Theoretical Analysis of Grokking Modular Addition
by: Mohamadi, Mohamad Amin, et al.
Published: (2024)
by: Mohamadi, Mohamad Amin, et al.
Published: (2024)
Clustering and Alignment: Understanding the Training Dynamics in Modular Addition
by: Musat, Tiberiu
Published: (2024)
by: Musat, Tiberiu
Published: (2024)
Learning Large-Scale Modular Addition with an Auxiliary Modulus
by: Kikuchi, Hanato, et al.
Published: (2026)
by: Kikuchi, Hanato, et al.
Published: (2026)
KAN: Kolmogorov-Arnold Networks
by: Liu, Ziming, et al.
Published: (2024)
by: Liu, Ziming, et al.
Published: (2024)
Understanding sparse autoencoder scaling in the presence of feature manifolds
by: Michaud, Eric J., et al.
Published: (2025)
by: Michaud, Eric J., et al.
Published: (2025)
The Remarkable Robustness of LLMs: Stages of Inference?
by: Lad, Vedang, et al.
Published: (2024)
by: Lad, Vedang, et al.
Published: (2024)
Scaling Laws For Scalable Oversight
by: Engels, Joshua, et al.
Published: (2025)
by: Engels, Joshua, et al.
Published: (2025)
ARC Is a Vision Problem!
by: Hu, Keya, et al.
Published: (2025)
by: Hu, Keya, et al.
Published: (2025)
From Self-Fulfilment to Survival of the Fittest
by: Mazierska, Ewa
Published: (2020)
by: Mazierska, Ewa
Published: (2020)
(How) Do Language Models Track State?
by: Li, Belinda Z., et al.
Published: (2025)
by: Li, Belinda Z., et al.
Published: (2025)
Efficient Hyperdimensional Computing with Modular Composite Representations
by: Angioli, Marco, et al.
Published: (2025)
by: Angioli, Marco, et al.
Published: (2025)
Uncovering a Universal Abstract Algorithm for Modular Addition in Neural Networks
by: McCracken, Gavin, et al.
Published: (2025)
by: McCracken, Gavin, et al.
Published: (2025)
Similar Items
-
The Quantization Model of Neural Scaling
by: Michaud, Eric J., et al.
Published: (2023) -
Physics of Skill Learning
by: Liu, Ziming, et al.
Published: (2025) -
On the creation of narrow AI: hierarchy and nonlocality of neural network skills
by: Michaud, Eric J., et al.
Published: (2025) -
A Neural Scaling Law from Lottery Ticket Ensembling
by: Liu, Ziming, et al.
Published: (2023) -
Opening the AI black box: program synthesis via mechanistic interpretability
by: Michaud, Eric J., et al.
Published: (2024)