Automated Interpretability Metrics Do Not Distinguish Trained and Random Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Heap, Thomas, Lawson, Tim, Farnik, Lucy, Aitchison, Laurence |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Residual Stream Analysis with Multi-Layer SAEs
by: Lawson, Tim, et al.
Published: (2024)
by: Lawson, Tim, et al.
Published: (2024)
Jacobian Sparse Autoencoders: Sparsify Computations, Not Just Activations
by: Farnik, Lucy, et al.
Published: (2025)
by: Farnik, Lucy, et al.
Published: (2025)
Learning to Skip the Middle Layers of Transformers
by: Lawson, Tim, et al.
Published: (2025)
by: Lawson, Tim, et al.
Published: (2025)
Massively Parallel Expectation Maximization For Approximate Posteriors
by: Heap, Thomas, et al.
Published: (2025)
by: Heap, Thomas, et al.
Published: (2025)
Why you don't overfit, and don't need Bayes if you only train for one epoch
by: Aitchison, Laurence
Published: (2024)
by: Aitchison, Laurence
Published: (2024)
Controlling changes to attention logits
by: Anson, Ben, et al.
Published: (2025)
by: Anson, Ben, et al.
Published: (2025)
Batch size invariant Adam
by: Wang, Xi, et al.
Published: (2024)
by: Wang, Xi, et al.
Published: (2024)
How to set AdamW's weight decay as you scale model and dataset size
by: Wang, Xi, et al.
Published: (2024)
by: Wang, Xi, et al.
Published: (2024)
Using Autodiff to Estimate Posterior Moments, Marginals and Samples
by: Bowyer, Sam, et al.
Published: (2023)
by: Bowyer, Sam, et al.
Published: (2023)
Function-Space Learning Rates
by: Milsom, Edward, et al.
Published: (2025)
by: Milsom, Edward, et al.
Published: (2025)
Using Neural Networks for Data Cleaning in Weather Datasets
by: Hanslope, Jack R. P., et al.
Published: (2024)
by: Hanslope, Jack R. P., et al.
Published: (2024)
Flexible Infinite-Width Graph Convolutional Neural Networks
by: Anson, Ben, et al.
Published: (2024)
by: Anson, Ben, et al.
Published: (2024)
Convolutional Deep Kernel Machines
by: Milsom, Edward, et al.
Published: (2023)
by: Milsom, Edward, et al.
Published: (2023)
Stochastic Kernel Regularisation Improves Generalisation in Deep Kernel Machines
by: Milsom, Edward, et al.
Published: (2024)
by: Milsom, Edward, et al.
Published: (2024)
Scale-invariant Attention
by: Anson, Ben, et al.
Published: (2025)
by: Anson, Ben, et al.
Published: (2025)
MONGOOSE: Path-wise Smooth Bayesian Optimisation via Meta-learning
by: Yang, Adam X., et al.
Published: (2023)
by: Yang, Adam X., et al.
Published: (2023)
Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints
by: Bowyer, Sam, et al.
Published: (2025)
by: Bowyer, Sam, et al.
Published: (2025)
Bayesian Low-rank Adaptation for Large Language Models
by: Yang, Adam X., et al.
Published: (2023)
by: Yang, Adam X., et al.
Published: (2023)
Interpretable Clustering with the Distinguishability Criterion
by: Turfah, Ali, et al.
Published: (2024)
by: Turfah, Ali, et al.
Published: (2024)
Inverse-Free Sparse Variational Gaussian Processes
by: Cortinovis, Stefano, et al.
Published: (2026)
by: Cortinovis, Stefano, et al.
Published: (2026)
Rodent-Bench
by: Heap, Thomas, et al.
Published: (2026)
by: Heap, Thomas, et al.
Published: (2026)
STARC: A General Framework For Quantifying Differences Between Reward Functions
by: Skalse, Joar, et al.
Published: (2023)
by: Skalse, Joar, et al.
Published: (2023)
Leveraging Interpretability in the Transformer to Automate the Proactive Scaling of Cloud Resources
by: Ba, Amadou, et al.
Published: (2024)
by: Ba, Amadou, et al.
Published: (2024)
Learning Generation Orders for Masked Discrete Diffusion Models via Variational Inference
by: Fox, David, et al.
Published: (2026)
by: Fox, David, et al.
Published: (2026)
Bayesian Reward Models for LLM Alignment
by: Yang, Adam X., et al.
Published: (2024)
by: Yang, Adam X., et al.
Published: (2024)
Questionable practices in machine learning
by: Leech, Gavin, et al.
Published: (2024)
by: Leech, Gavin, et al.
Published: (2024)
Graph Homomorphism Distortion: A Metric to Distinguish Them All and in the Latent Space Bind Them
by: Carrasco, Martin, et al.
Published: (2025)
by: Carrasco, Martin, et al.
Published: (2025)
Prior Aware Memorization: An Efficient Metric for Distinguishing Memorization from Generalization in Large Language Models
by: Tiwari, Trishita, et al.
Published: (2026)
by: Tiwari, Trishita, et al.
Published: (2026)
Inducing Human-like Biases in Moral Reasoning Language Models
by: Karpov, Artem, et al.
Published: (2024)
by: Karpov, Artem, et al.
Published: (2024)
Surrogate Fitness Metrics for Interpretable Reinforcement Learning
by: Altmann, Philipp, et al.
Published: (2025)
by: Altmann, Philipp, et al.
Published: (2025)
Symmetry-Aware Transformer Training for Automated Planning
by: Fritzsche, Markus, et al.
Published: (2025)
by: Fritzsche, Markus, et al.
Published: (2025)
Structured Transformations for Stable and Interpretable Neural Computation
by: Nikooroo, Saleh, et al.
Published: (2025)
by: Nikooroo, Saleh, et al.
Published: (2025)
Machine learning emulation of precipitation from km-scale UK regional climate simulations using a diffusion model
by: Addison, Henry, et al.
Published: (2024)
by: Addison, Henry, et al.
Published: (2024)
Metric as Transform: Exploring beyond Affine Transform for Interpretable Neural Network
by: Sapkota, Suman
Published: (2024)
by: Sapkota, Suman
Published: (2024)
Symmetry Breaking in Transformers for Efficient and Interpretable Training
by: Silverstein, Eva, et al.
Published: (2026)
by: Silverstein, Eva, et al.
Published: (2026)
Random-Effects Algorithm for Random Objects in Metric Spaces
by: Matabuena, Marcos, et al.
Published: (2026)
by: Matabuena, Marcos, et al.
Published: (2026)
When Are Two Networks the Same? Tensor Similarity for Mechanistic Interpretability
by: Gonzalez, ML Nissen, et al.
Published: (2026)
by: Gonzalez, ML Nissen, et al.
Published: (2026)
Automated Classification of Volcanic Earthquakes Using Transformer Encoders: Insights into Data Quality and Model Interpretability
by: Suzuki, Y., et al.
Published: (2025)
by: Suzuki, Y., et al.
Published: (2025)
Deterministic Bounds and Random Estimates of Metric Tensors on Neuromanifolds
by: Sun, Ke
Published: (2025)
by: Sun, Ke
Published: (2025)
ML Interpretability: Simple Isn't Easy
by: Räz, Tim
Published: (2022)
by: Räz, Tim
Published: (2022)
Similar Items
-
Residual Stream Analysis with Multi-Layer SAEs
by: Lawson, Tim, et al.
Published: (2024) -
Jacobian Sparse Autoencoders: Sparsify Computations, Not Just Activations
by: Farnik, Lucy, et al.
Published: (2025) -
Learning to Skip the Middle Layers of Transformers
by: Lawson, Tim, et al.
Published: (2025) -
Massively Parallel Expectation Maximization For Approximate Posteriors
by: Heap, Thomas, et al.
Published: (2025) -
Why you don't overfit, and don't need Bayes if you only train for one epoch
by: Aitchison, Laurence
Published: (2024)