Evaluating SAE interpretability without explanations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Paulo, Gonçalo, Belrose, Nora |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Sparse Autoencoders Trained on the Same Data Learn Different Features
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2025)
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2025)
Partially Rewriting a Transformer in Natural Language
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2025)
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2025)
Transcoders Beat Sparse Autoencoders for Interpretability
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2025)
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2025)
Does Transformer Interpretability Transfer to RNNs?
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2024)
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2024)
Automatically Interpreting Millions of Features in Large Language Models
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2024)
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2024)
Estimating the Probability of Sampling a Trained Neural Network at Random
von: Scherlis, Adam, et al.
Veröffentlicht: (2025)
von: Scherlis, Adam, et al.
Veröffentlicht: (2025)
Slowing Learning by Erasing Simple Features
von: Quirke, Lucia, et al.
Veröffentlicht: (2025)
von: Quirke, Lucia, et al.
Veröffentlicht: (2025)
Converting MLPs into Polynomials in Closed Form
von: Belrose, Nora, et al.
Veröffentlicht: (2025)
von: Belrose, Nora, et al.
Veröffentlicht: (2025)
Balancing Label Quantity and Quality for Scalable Elicitation
von: Mallen, Alex, et al.
Veröffentlicht: (2024)
von: Mallen, Alex, et al.
Veröffentlicht: (2024)
Understanding Gradient Descent through the Training Jacobian
von: Belrose, Nora, et al.
Veröffentlicht: (2024)
von: Belrose, Nora, et al.
Veröffentlicht: (2024)
Examining Two Hop Reasoning Through Information Content Scaling
von: Johnston, David, et al.
Veröffentlicht: (2025)
von: Johnston, David, et al.
Veröffentlicht: (2025)
Binary Sparse Coding for Interpretability
von: Quirke, Lucia, et al.
Veröffentlicht: (2025)
von: Quirke, Lucia, et al.
Veröffentlicht: (2025)
Refusal in LLMs is an Affine Function
von: Marshall, Thomas, et al.
Veröffentlicht: (2024)
von: Marshall, Thomas, et al.
Veröffentlicht: (2024)
Evaluating Synthetic Activations composed of SAE Latents in GPT-2
von: Giglemiani, Giorgi, et al.
Veröffentlicht: (2024)
von: Giglemiani, Giorgi, et al.
Veröffentlicht: (2024)
Mechanistic Anomaly Detection for "Quirky" Language Models
von: Johnston, David O., et al.
Veröffentlicht: (2025)
von: Johnston, David O., et al.
Veröffentlicht: (2025)
Neural Networks Learn Statistics of Increasing Complexity
von: Belrose, Nora, et al.
Veröffentlicht: (2024)
von: Belrose, Nora, et al.
Veröffentlicht: (2024)
Eliciting Latent Knowledge from Quirky Language Models
von: Mallen, Alex, et al.
Veröffentlicht: (2023)
von: Mallen, Alex, et al.
Veröffentlicht: (2023)
Symmetric observations without symmetric causal explanations
von: William, Christian, et al.
Veröffentlicht: (2025)
von: William, Christian, et al.
Veröffentlicht: (2025)
Tokenized SAEs: Disentangling SAE Reconstructions
von: Dooms, Thomas, et al.
Veröffentlicht: (2025)
von: Dooms, Thomas, et al.
Veröffentlicht: (2025)
LEACE: Perfect linear concept erasure in closed form
von: Belrose, Nora, et al.
Veröffentlicht: (2023)
von: Belrose, Nora, et al.
Veröffentlicht: (2023)
Eliciting Latent Predictions from Transformers with the Tuned Lens
von: Belrose, Nora, et al.
Veröffentlicht: (2023)
von: Belrose, Nora, et al.
Veröffentlicht: (2023)
Neuron-based explanations of neural networks sacrifice completeness and interpretability
von: Dey, Nolan, et al.
Veröffentlicht: (2020)
von: Dey, Nolan, et al.
Veröffentlicht: (2020)
Modular addition without black-boxes: Compressing explanations of MLPs that compute numerical integration
von: Yip, Chun Hei, et al.
Veröffentlicht: (2024)
von: Yip, Chun Hei, et al.
Veröffentlicht: (2024)
Evolution of SAE Features Across Layers in LLMs
von: Balcells, Daniel, et al.
Veröffentlicht: (2024)
von: Balcells, Daniel, et al.
Veröffentlicht: (2024)
SAE: Single Architecture Ensemble Neural Networks
von: Ferianc, Martin, et al.
Veröffentlicht: (2024)
von: Ferianc, Martin, et al.
Veröffentlicht: (2024)
FaithfulSAE: Towards Capturing Faithful Features with Sparse Autoencoders without External Dataset Dependencies
von: Cho, Seonglae, et al.
Veröffentlicht: (2025)
von: Cho, Seonglae, et al.
Veröffentlicht: (2025)
Evaluating the stability of model explanations in instance-dependent cost-sensitive credit scoring
von: Ballegeer, Matteo, et al.
Veröffentlicht: (2025)
von: Ballegeer, Matteo, et al.
Veröffentlicht: (2025)
On GNN explanability with activation rules
von: Veyrin-Forrer, Luca, et al.
Veröffentlicht: (2024)
von: Veyrin-Forrer, Luca, et al.
Veröffentlicht: (2024)
The effect of whitening on explanation performance
von: Clark, Benedict, et al.
Veröffentlicht: (2026)
von: Clark, Benedict, et al.
Veröffentlicht: (2026)
SCALAR: Benchmarking SAE Interaction Sparsity in Toy LLMs
von: Fillingham, Sean P., et al.
Veröffentlicht: (2025)
von: Fillingham, Sean P., et al.
Veröffentlicht: (2025)
Mathematically rigorous proofs for Shapley explanations
von: van Batenburg, David
Veröffentlicht: (2025)
von: van Batenburg, David
Veröffentlicht: (2025)
OrtSAE: Orthogonal Sparse Autoencoders Uncover Atomic Features
von: Korznikov, Anton, et al.
Veröffentlicht: (2025)
von: Korznikov, Anton, et al.
Veröffentlicht: (2025)
Concept-SAE: Active Causal Probing of Visual Model Behavior
von: Ding, Jianrong, et al.
Veröffentlicht: (2025)
von: Ding, Jianrong, et al.
Veröffentlicht: (2025)
Tree SAE: Learning Hierarchical Feature Structures in Sparse Autoencoders
von: Cao, Tue M., et al.
Veröffentlicht: (2026)
von: Cao, Tue M., et al.
Veröffentlicht: (2026)
AlignSAE: Concept-Aligned Sparse Autoencoders
von: Yang, Minglai, et al.
Veröffentlicht: (2025)
von: Yang, Minglai, et al.
Veröffentlicht: (2025)
The explanation dialogues: an expert focus study to understand requirements towards explanations within the GDPR
von: State, Laura, et al.
Veröffentlicht: (2025)
von: State, Laura, et al.
Veröffentlicht: (2025)
Effector: A Python package for regional explanations
von: Gkolemis, Vasilis, et al.
Veröffentlicht: (2024)
von: Gkolemis, Vasilis, et al.
Veröffentlicht: (2024)
Dense SAE Latents Are Features, Not Bugs
von: Sun, Xiaoqing, et al.
Veröffentlicht: (2025)
von: Sun, Xiaoqing, et al.
Veröffentlicht: (2025)
Compositional Literary Primitives in Instruction-Tuned LLMs: Cross-Architectural SAE Features for Self, Style, and Affect
von: Presa, Joao Paulo Cavalcante, et al.
Veröffentlicht: (2026)
von: Presa, Joao Paulo Cavalcante, et al.
Veröffentlicht: (2026)
Integrating attention into explanation frameworks for language and vision transformers
von: Eggen, Marte, et al.
Veröffentlicht: (2025)
von: Eggen, Marte, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Sparse Autoencoders Trained on the Same Data Learn Different Features
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2025) -
Partially Rewriting a Transformer in Natural Language
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2025) -
Transcoders Beat Sparse Autoencoders for Interpretability
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2025) -
Does Transformer Interpretability Transfer to RNNs?
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2024) -
Automatically Interpreting Millions of Features in Large Language Models
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2024)