Interpretability as Compression: Reconsidering SAE Explanations of Neural Activations with MDL-SAEs
Fuente:
arXiv
Saved in:
| Main Authors: | Ayonrinde, Kola, Pearce, Michael T., Sharkey, Lee |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Mathematical Philosophy of Explanations in Mechanistic Interpretability -- The Strange Science Part I.i
by: Ayonrinde, Kola, et al.
Published: (2025)
by: Ayonrinde, Kola, et al.
Published: (2025)
Adaptive Sparse Allocation with Mutual Choice & Feature Choice Sparse Autoencoders
by: Ayonrinde, Kola
Published: (2024)
by: Ayonrinde, Kola
Published: (2024)
Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii
by: Ayonrinde, Kola, et al.
Published: (2025)
by: Ayonrinde, Kola, et al.
Published: (2025)
AlphaZip: Neural Network-Enhanced Lossless Text Compression
by: Narashiman, Swathi Shree, et al.
Published: (2024)
by: Narashiman, Swathi Shree, et al.
Published: (2024)
FibQuant: Universal Vector Quantization for Random-Access KV-Cache Compression
by: Lee, Namyoon, et al.
Published: (2026)
by: Lee, Namyoon, et al.
Published: (2026)
Robust Counterfactual Explanations for Neural Networks With Probabilistic Guarantees
by: Hamman, Faisal, et al.
Published: (2023)
by: Hamman, Faisal, et al.
Published: (2023)
Optimizing Learned Image Compression on Scalar and Entropy-Constraint Quantization
by: Borzechowski, Florian, et al.
Published: (2025)
by: Borzechowski, Florian, et al.
Published: (2025)
Haiku to Opus in Just 10 bits: LLMs Unlock Massive Compression Gains
by: Rinberg, Roy, et al.
Published: (2026)
by: Rinberg, Roy, et al.
Published: (2026)
Distributed and Rate-Adaptive Feature Compression
by: Deshmukh, Aditya, et al.
Published: (2024)
by: Deshmukh, Aditya, et al.
Published: (2024)
Compressing Chemistry Reveals Functional Groups
by: Sharma, Ruben, et al.
Published: (2025)
by: Sharma, Ruben, et al.
Published: (2025)
Interpretable Diffusion via Information Decomposition
by: Kong, Xianghao, et al.
Published: (2023)
by: Kong, Xianghao, et al.
Published: (2023)
Partial Information Decomposition for Data Interpretability and Feature Selection
by: Westphal, Charles, et al.
Published: (2024)
by: Westphal, Charles, et al.
Published: (2024)
A General Error-Theoretical Analysis Framework for Constructing Compression Strategies
by: Zhang, Boyang, et al.
Published: (2025)
by: Zhang, Boyang, et al.
Published: (2025)
Flexible Variational Information Bottleneck: Achieving Diverse Compression with a Single Training
by: Kudo, Sota, et al.
Published: (2024)
by: Kudo, Sota, et al.
Published: (2024)
GeoIB: Geometry-Aware Information Bottleneck via Statistical-Manifold Compression
by: Wang, Weiqi, et al.
Published: (2026)
by: Wang, Weiqi, et al.
Published: (2026)
Broadcast Channel Cooperative Gain: An Operational Interpretation of Partial Information Decomposition
by: Tian, Chao, et al.
Published: (2025)
by: Tian, Chao, et al.
Published: (2025)
Compression via Pre-trained Transformers: A Study on Byte-Level Multimodal Data
by: Heurtel-Depeiges, David, et al.
Published: (2024)
by: Heurtel-Depeiges, David, et al.
Published: (2024)
Understanding LLM Behaviors via Compression: Data Generation, Knowledge Acquisition and Scaling Laws
by: Pan, Zhixuan, et al.
Published: (2025)
by: Pan, Zhixuan, et al.
Published: (2025)
PAGE: Prototype-Based Model-Level Explanations for Graph Neural Networks
by: Shin, Yong-Min, et al.
Published: (2022)
by: Shin, Yong-Min, et al.
Published: (2022)
Neural Networks Learn Generic Multi-Index Models Near Information-Theoretic Limit
by: Zhang, Bohan, et al.
Published: (2025)
by: Zhang, Bohan, et al.
Published: (2025)
Informationally Compressive Anonymization: Non-Degrading Sensitive Input Protection for Privacy-Preserving Supervised Machine Learning
by: Samuelson, Jeremy J
Published: (2026)
by: Samuelson, Jeremy J
Published: (2026)
SPEX: Scaling Feature Interaction Explanations for LLMs
by: Kang, Justin Singh, et al.
Published: (2025)
by: Kang, Justin Singh, et al.
Published: (2025)
Polynomial Context-Truncation Sensitivity in Autoregressive Language Models: Sequential Wyner-Ziv Bounds for KV Cache Compression
by: Kim, Munsik
Published: (2026)
by: Kim, Munsik
Published: (2026)
Neural Polar Decoders for Deletion Channels
by: Aharoni, Ziv, et al.
Published: (2025)
by: Aharoni, Ziv, et al.
Published: (2025)
Language Modeling Is Compression
by: Delétang, Grégoire, et al.
Published: (2023)
by: Delétang, Grégoire, et al.
Published: (2023)
MDL: A Unified Multi-Distribution Learner in Large-scale Industrial Recommendation through Tokenization
by: Mu, Shanlei, et al.
Published: (2026)
by: Mu, Shanlei, et al.
Published: (2026)
Learning in Convolutional Neural Networks Accelerated by Transfer Entropy
by: Moldovan, Adrian, et al.
Published: (2024)
by: Moldovan, Adrian, et al.
Published: (2024)
Neural Beam Field for Spatial Beam RSRP Prediction
by: Guo, Keqiang, et al.
Published: (2025)
by: Guo, Keqiang, et al.
Published: (2025)
Lost and Found in Translation: Variational Diagnostics for Neural Codebook Channels
by: Hayashi, Yusuke
Published: (2026)
by: Hayashi, Yusuke
Published: (2026)
Compression Represents Intelligence Linearly
by: Huang, Yuzhen, et al.
Published: (2024)
by: Huang, Yuzhen, et al.
Published: (2024)
Can Kernel Methods Explain How the Data Affects Neural Collapse?
by: Kothapalli, Vignesh, et al.
Published: (2024)
by: Kothapalli, Vignesh, et al.
Published: (2024)
Neural Estimation of Pairwise Mutual Information in Masked Discrete Sequence Models
by: Sharma, Jai, et al.
Published: (2026)
by: Sharma, Jai, et al.
Published: (2026)
Analyzing Shapley Additive Explanations to Understand Anomaly Detection Algorithm Behaviors and Their Complementarity
by: Levy, Jordan, et al.
Published: (2026)
by: Levy, Jordan, et al.
Published: (2026)
Enhancing Explainability of Graph Neural Networks Through Conceptual and Structural Analyses and Their Extensions
by: Bui, Tien Cuong
Published: (2025)
by: Bui, Tien Cuong
Published: (2025)
Learning During Detection: Continual Learning for Neural OFDM Receivers via DMRS
by: Obeed, Mohanad, et al.
Published: (2026)
by: Obeed, Mohanad, et al.
Published: (2026)
Deep Randomized Distributed Function Computation (DeepRDFC): Neural Distributed Channel Simulation
by: Bergström, Didrik, et al.
Published: (2026)
by: Bergström, Didrik, et al.
Published: (2026)
Memorization-Compression Cycles Improve Generalization
by: Yu, Fangyuan
Published: (2025)
by: Yu, Fangyuan
Published: (2025)
Neural Channel Knowledge Map Assisted Scheduling Optimization of Active IRSs in Multi-User Systems
by: Chen, Xintong, et al.
Published: (2025)
by: Chen, Xintong, et al.
Published: (2025)
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs
by: Song, Xiangchen, et al.
Published: (2025)
by: Song, Xiangchen, et al.
Published: (2025)
LASER: Linear Compression in Wireless Distributed Optimization
by: Makkuva, Ashok Vardhan, et al.
Published: (2023)
by: Makkuva, Ashok Vardhan, et al.
Published: (2023)
Similar Items
-
A Mathematical Philosophy of Explanations in Mechanistic Interpretability -- The Strange Science Part I.i
by: Ayonrinde, Kola, et al.
Published: (2025) -
Adaptive Sparse Allocation with Mutual Choice & Feature Choice Sparse Autoencoders
by: Ayonrinde, Kola
Published: (2024) -
Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii
by: Ayonrinde, Kola, et al.
Published: (2025) -
AlphaZip: Neural Network-Enhanced Lossless Text Compression
by: Narashiman, Swathi Shree, et al.
Published: (2024) -
FibQuant: Universal Vector Quantization for Random-Access KV-Cache Compression
by: Lee, Namyoon, et al.
Published: (2026)