Compute Optimal Inference and Provable Amortisation Gap in Sparse Autoencoders
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | O'Neill, Charles, Gumran, Alim, Klindt, David |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Sparse Autoencoders Enable Scalable and Reliable Circuit Identification in Language Models
von: O'Neill, Charles, et al.
Veröffentlicht: (2024)
von: O'Neill, Charles, et al.
Veröffentlicht: (2024)
Resurrecting the Salmon: Rethinking Mechanistic Interpretability with Domain-Specific Sparse Autoencoders
von: O'Neill, Charles, et al.
Veröffentlicht: (2025)
von: O'Neill, Charles, et al.
Veröffentlicht: (2025)
Disentangling Dense Embeddings with Sparse Autoencoders
von: O'Neill, Charles, et al.
Veröffentlicht: (2024)
von: O'Neill, Charles, et al.
Veröffentlicht: (2024)
Data Whitening Improves Sparse Autoencoder Learning
von: Saraswatula, Ashwin, et al.
Veröffentlicht: (2025)
von: Saraswatula, Ashwin, et al.
Veröffentlicht: (2025)
From superposition to sparse codes: interpretable representations in neural networks
von: Klindt, David, et al.
Veröffentlicht: (2025)
von: Klindt, David, et al.
Veröffentlicht: (2025)
Self-Attention as a Parametric Endofunctor: A Categorical Framework for Transformer Architectures
von: O'Neill, Charles
Veröffentlicht: (2025)
von: O'Neill, Charles
Veröffentlicht: (2025)
Amortising Inference and Meta-Learning Priors in Neural Networks
von: Rochussen, Tommy, et al.
Veröffentlicht: (2026)
von: Rochussen, Tommy, et al.
Veröffentlicht: (2026)
Stop Probing, Start Coding: Why Linear Probes and Sparse Autoencoders Fail at Compositional Generalisation
von: Pacela, Vitória Barin, et al.
Veröffentlicht: (2026)
von: Pacela, Vitória Barin, et al.
Veröffentlicht: (2026)
Type 2 Tobit Sample Selection Models with Bayesian Additive Regression Trees
von: O'Neill, Eoghan
Veröffentlicht: (2025)
von: O'Neill, Eoghan
Veröffentlicht: (2025)
Grokking Beyond Neural Networks: An Empirical Exploration with Model Complexity
von: Miller, Jack, et al.
Veröffentlicht: (2023)
von: Miller, Jack, et al.
Veröffentlicht: (2023)
Amortised and provably-robust simulation-based inference
von: Bharti, Ayush, et al.
Veröffentlicht: (2026)
von: Bharti, Ayush, et al.
Veröffentlicht: (2026)
Towards Interpretable and Inference-Optimal COT Reasoning with Sparse Autoencoder-Guided Generation
von: Zhao, Daniel, et al.
Veröffentlicht: (2025)
von: Zhao, Daniel, et al.
Veröffentlicht: (2025)
Sketching the Heat Kernel: Using Gaussian Processes to Embed Data
von: Gilbert, Anna C., et al.
Veröffentlicht: (2024)
von: Gilbert, Anna C., et al.
Veröffentlicht: (2024)
Superposition disentanglement of neural representations reveals hidden alignment
von: Longon, André, et al.
Veröffentlicht: (2025)
von: Longon, André, et al.
Veröffentlicht: (2025)
Measuring Sharpness in Grokking
von: Miller, Jack, et al.
Veröffentlicht: (2024)
von: Miller, Jack, et al.
Veröffentlicht: (2024)
Modelling the Doughnut of social and planetary boundaries with frugal machine learning
von: Vrizzi, Stefano, et al.
Veröffentlicht: (2025)
von: Vrizzi, Stefano, et al.
Veröffentlicht: (2025)
Taming Polysemanticity in LLMs: Provable Feature Recovery via Sparse Autoencoders
von: Chen, Siyu, et al.
Veröffentlicht: (2025)
von: Chen, Siyu, et al.
Veröffentlicht: (2025)
Are Sparse Autoencoder Benchmarks Reliable?
von: Chanin, David
Veröffentlicht: (2026)
von: Chanin, David
Veröffentlicht: (2026)
A Single Direction of Truth: An Observer Model's Linear Residual Probe Exposes and Steers Contextual Hallucinations
von: O'Neill, Charles, et al.
Veröffentlicht: (2025)
von: O'Neill, Charles, et al.
Veröffentlicht: (2025)
Sparse Autoencoders, Again?
von: Lu, Yin, et al.
Veröffentlicht: (2025)
von: Lu, Yin, et al.
Veröffentlicht: (2025)
Occam's Razor for Self Supervised Learning: What is Sufficient to Learn Good Representations?
von: Ibrahim, Mark, et al.
Veröffentlicht: (2024)
von: Ibrahim, Mark, et al.
Veröffentlicht: (2024)
Compute-Optimal LLMs Provably Generalize Better With Scale
von: Finzi, Marc, et al.
Veröffentlicht: (2025)
von: Finzi, Marc, et al.
Veröffentlicht: (2025)
Steering Language Model Refusal with Sparse Autoencoders
von: O'Brien, Kyle, et al.
Veröffentlicht: (2024)
von: O'Brien, Kyle, et al.
Veröffentlicht: (2024)
Statistical Inference in Tensor Completion: Optimal Uncertainty Quantification and Statistical-to-Computational Gaps
von: Ma, Wanteng, et al.
Veröffentlicht: (2024)
von: Ma, Wanteng, et al.
Veröffentlicht: (2024)
Low-Rank Key Value Attention
von: O'Neill, James, et al.
Veröffentlicht: (2026)
von: O'Neill, James, et al.
Veröffentlicht: (2026)
Ensembling Sparse Autoencoders
von: Gadgil, Soham, et al.
Veröffentlicht: (2025)
von: Gadgil, Soham, et al.
Veröffentlicht: (2025)
CA-PCA: Manifold Dimension Estimation, Adapted for Curvature
von: Gilbert, Anna C., et al.
Veröffentlicht: (2023)
von: Gilbert, Anna C., et al.
Veröffentlicht: (2023)
Semantic Optimal Transport for Sparse Autoencoder Feature Matching and Circuit Compression
von: Cao, Tue M., et al.
Veröffentlicht: (2026)
von: Cao, Tue M., et al.
Veröffentlicht: (2026)
Toward Identifiable Sparse Autoencoders
von: Nelson, Walter, et al.
Veröffentlicht: (2026)
von: Nelson, Walter, et al.
Veröffentlicht: (2026)
Analysis of Variational Sparse Autoencoders
von: Baker, Zachary, et al.
Veröffentlicht: (2025)
von: Baker, Zachary, et al.
Veröffentlicht: (2025)
Jacobian Sparse Autoencoders: Sparsify Computations, Not Just Activations
von: Farnik, Lucy, et al.
Veröffentlicht: (2025)
von: Farnik, Lucy, et al.
Veröffentlicht: (2025)
Transformers Provably Learn Sparse XOR with Polylogarithmic Parameters
von: Han, Yaomengxi, et al.
Veröffentlicht: (2025)
von: Han, Yaomengxi, et al.
Veröffentlicht: (2025)
Enhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders
von: Marks, Luke, et al.
Veröffentlicht: (2024)
von: Marks, Luke, et al.
Veröffentlicht: (2024)
Compression of Structured Data with Autoencoders: Provable Benefit of Nonlinearities and Depth
von: Kögler, Kevin, et al.
Veröffentlicht: (2024)
von: Kögler, Kevin, et al.
Veröffentlicht: (2024)
Position: An Empirically Grounded Identifiability Theory Will Accelerate Self-Supervised Learning Research
von: Reizinger, Patrik, et al.
Veröffentlicht: (2025)
von: Reizinger, Patrik, et al.
Veröffentlicht: (2025)
Decomposing The Dark Matter of Sparse Autoencoders
von: Engels, Joshua, et al.
Veröffentlicht: (2024)
von: Engels, Joshua, et al.
Veröffentlicht: (2024)
Transcoders Beat Sparse Autoencoders for Interpretability
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2025)
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2025)
Evaluating Sparse Autoencoders for Monosemantic Representation
von: Fereidouni, Moghis, et al.
Veröffentlicht: (2025)
von: Fereidouni, Moghis, et al.
Veröffentlicht: (2025)
Inference via Interpolation: Contrastive Representations Provably Enable Planning and Inference
von: Eysenbach, Benjamin, et al.
Veröffentlicht: (2024)
von: Eysenbach, Benjamin, et al.
Veröffentlicht: (2024)
Efficient Dictionary Learning with Switch Sparse Autoencoders
von: Mudide, Anish, et al.
Veröffentlicht: (2024)
von: Mudide, Anish, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Sparse Autoencoders Enable Scalable and Reliable Circuit Identification in Language Models
von: O'Neill, Charles, et al.
Veröffentlicht: (2024) -
Resurrecting the Salmon: Rethinking Mechanistic Interpretability with Domain-Specific Sparse Autoencoders
von: O'Neill, Charles, et al.
Veröffentlicht: (2025) -
Disentangling Dense Embeddings with Sparse Autoencoders
von: O'Neill, Charles, et al.
Veröffentlicht: (2024) -
Data Whitening Improves Sparse Autoencoder Learning
von: Saraswatula, Ashwin, et al.
Veröffentlicht: (2025) -
From superposition to sparse codes: interpretable representations in neural networks
von: Klindt, David, et al.
Veröffentlicht: (2025)