Opening the AI black box: program synthesis via mechanistic interpretability
Fuente:
arXiv
Saved in:
| Main Authors: | Michaud, Eric J., Liao, Isaac, Lad, Vedang, Liu, Ziming, Mudide, Anish, Loughridge, Chloe, Guo, Zifan Carl, Kheirkhah, Tara Rezaei, Vukelić, Mateja, Tegmark, Max |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient Dictionary Learning with Switch Sparse Autoencoders
by: Mudide, Anish, et al.
Published: (2024)
by: Mudide, Anish, et al.
Published: (2024)
The Remarkable Robustness of LLMs: Stages of Inference?
by: Lad, Vedang, et al.
Published: (2024)
by: Lad, Vedang, et al.
Published: (2024)
Opening the black box of language acquisition
by: Michaud, Jérôme, et al.
Published: (2024)
by: Michaud, Jérôme, et al.
Published: (2024)
Survival of the Fittest Representation: A Case Study with Modular Addition
by: Ding, Xiaoman Delores, et al.
Published: (2024)
by: Ding, Xiaoman Delores, et al.
Published: (2024)
DafnyBench: A Benchmark for Formal Software Verification
by: Loughridge, Chloe, et al.
Published: (2024)
by: Loughridge, Chloe, et al.
Published: (2024)
The Quantization Model of Neural Scaling
by: Michaud, Eric J., et al.
Published: (2023)
by: Michaud, Eric J., et al.
Published: (2023)
A Neural Scaling Law from Lottery Ticket Ensembling
by: Liu, Ziming, et al.
Published: (2023)
by: Liu, Ziming, et al.
Published: (2023)
Physics of Skill Learning
by: Liu, Ziming, et al.
Published: (2025)
by: Liu, Ziming, et al.
Published: (2025)
OptPDE: Discovering Novel Integrable Systems via AI-Human Collaboration
by: Kantamneni, Subhash, et al.
Published: (2024)
by: Kantamneni, Subhash, et al.
Published: (2024)
How Do Transformers "Do" Physics? Investigating the Simple Harmonic Oscillator
by: Kantamneni, Subhash, et al.
Published: (2024)
by: Kantamneni, Subhash, et al.
Published: (2024)
Do Two AI Scientists Agree?
by: Fu, Xinghong, et al.
Published: (2025)
by: Fu, Xinghong, et al.
Published: (2025)
Not All Language Model Features Are One-Dimensionally Linear
by: Engels, Joshua, et al.
Published: (2024)
by: Engels, Joshua, et al.
Published: (2024)
Universal Neurons in GPT2 Language Models
by: Gurnee, Wes, et al.
Published: (2024)
by: Gurnee, Wes, et al.
Published: (2024)
On the creation of narrow AI: hierarchy and nonlocality of neural network skills
by: Michaud, Eric J., et al.
Published: (2025)
by: Michaud, Eric J., et al.
Published: (2025)
GenEFT: Understanding Statics and Dynamics of Model Generalization via Effective Theory
by: Baek, David D., et al.
Published: (2024)
by: Baek, David D., et al.
Published: (2024)
A Resource Model For Neural Scaling Law
by: Song, Jinyeop, et al.
Published: (2024)
by: Song, Jinyeop, et al.
Published: (2024)
Neural Thermodynamic Laws for Large Language Model Training
by: Liu, Ziming, et al.
Published: (2025)
by: Liu, Ziming, et al.
Published: (2025)
dafny-annotator: AI-Assisted Verification of Dafny Programs
by: Poesia, Gabriel, et al.
Published: (2024)
by: Poesia, Gabriel, et al.
Published: (2024)
Library and Information Science Education in Morocco: Observations on a Recent Visit to the "Ecole des Sciences de l'Information" and the "Centre National de Documentation" in Rabat.
by: Loughridge, B.
Published: (1987)
by: Loughridge, B.
Published: (1987)
Harmonic Loss Trains Interpretable AI Models
by: Baek, David D., et al.
Published: (2025)
by: Baek, David D., et al.
Published: (2025)
A benchmark for vericoding: formally verified program synthesis
by: Bursuc, Sergiu, et al.
Published: (2025)
by: Bursuc, Sergiu, et al.
Published: (2025)
KAN 2.0: Kolmogorov-Arnold Networks Meet Science
by: Liu, Ziming, et al.
Published: (2024)
by: Liu, Ziming, et al.
Published: (2024)
Language Models Use Trigonometry to Do Addition
by: Kantamneni, Subhash, et al.
Published: (2025)
by: Kantamneni, Subhash, et al.
Published: (2025)
The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets
by: Marks, Samuel, et al.
Published: (2023)
by: Marks, Samuel, et al.
Published: (2023)
Language Models Represent Space and Time
by: Gurnee, Wes, et al.
Published: (2023)
by: Gurnee, Wes, et al.
Published: (2023)
Towards Understanding Distilled Reasoning Models: A Representational Approach
by: Baek, David D., et al.
Published: (2025)
by: Baek, David D., et al.
Published: (2025)
The Geometry of Concepts: Sparse Autoencoder Feature Structure
by: Li, Yuxiao, et al.
Published: (2024)
by: Li, Yuxiao, et al.
Published: (2024)
Surrogate modeling for interpreting black-box LLMs in medical predictions
by: Han, Changho, et al.
Published: (2026)
by: Han, Changho, et al.
Published: (2026)
Studying and measuring the amount of detergents in sea-water samples in order to eliminate with adsorbents and inorganic nano-adsorbents for the sake of securing the water for marine nourishment pools
by: Kheirkhah, Laya
Published: (2010)
by: Kheirkhah, Laya
Published: (2010)
Opening the black box: New insights into the role of temperature in the marine distributions of Pacific salmon
by: Joseph A. Langan, et al.
Published: (2024)
by: Joseph A. Langan, et al.
Published: (2024)
Opening the black box of recurrence in innovation policy. Sources of the Matthew effect in the Argentinean case
by: Diana Suarez
Published: (2021)
by: Diana Suarez
Published: (2021)
Opening the “black box”: Organizational differences between charter schools and traditional public schools
by: Xin Wei
Published: (2014)
by: Xin Wei
Published: (2014)
An introduction to graphical tensor notation for mechanistic interpretability
by: Taylor, Jordan K.
Published: (2024)
by: Taylor, Jordan K.
Published: (2024)
Causal interpretations can be based on mechanistic knowledge
by: James B. Grace, et al.
Published: (2025)
by: James B. Grace, et al.
Published: (2025)
Quantum circuit simulation of black hole evaporation and Maxwell demon interpretation
by: Hadi, H., et al.
Published: (2025)
by: Hadi, H., et al.
Published: (2025)
Opening Pandora's box of transport phenomena
by: Timo Vesala
Published: (2024)
by: Timo Vesala
Published: (2024)
Low-Rank Adapting Models for Sparse Autoencoders
by: Chen, Matthew, et al.
Published: (2025)
by: Chen, Matthew, et al.
Published: (2025)
Decomposing The Dark Matter of Sparse Autoencoders
by: Engels, Joshua, et al.
Published: (2024)
by: Engels, Joshua, et al.
Published: (2024)
A mechanistically interpretable neural network for regulatory genomics
by: Tseng, Alex M., et al.
Published: (2024)
by: Tseng, Alex M., et al.
Published: (2024)
Bilinear MLPs enable weight-based mechanistic interpretability
by: Pearce, Michael T., et al.
Published: (2024)
by: Pearce, Michael T., et al.
Published: (2024)
Similar Items
-
Efficient Dictionary Learning with Switch Sparse Autoencoders
by: Mudide, Anish, et al.
Published: (2024) -
The Remarkable Robustness of LLMs: Stages of Inference?
by: Lad, Vedang, et al.
Published: (2024) -
Opening the black box of language acquisition
by: Michaud, Jérôme, et al.
Published: (2024) -
Survival of the Fittest Representation: A Case Study with Modular Addition
by: Ding, Xiaoman Delores, et al.
Published: (2024) -
DafnyBench: A Benchmark for Formal Software Verification
by: Loughridge, Chloe, et al.
Published: (2024)