Binary Sparse Coding for Interpretability
Fuente:
arXiv
Salvato in:
| Autori principali: | Quirke, Lucia, Shabalin, Stepan, Belrose, Nora |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Transcoders Beat Sparse Autoencoders for Interpretability
di: Paulo, Gonçalo, et al.
Pubblicazione: (2025)
di: Paulo, Gonçalo, et al.
Pubblicazione: (2025)
Slowing Learning by Erasing Simple Features
di: Quirke, Lucia, et al.
Pubblicazione: (2025)
di: Quirke, Lucia, et al.
Pubblicazione: (2025)
Neural Networks Learn Statistics of Increasing Complexity
di: Belrose, Nora, et al.
Pubblicazione: (2024)
di: Belrose, Nora, et al.
Pubblicazione: (2024)
Sparse Autoencoders Trained on the Same Data Learn Different Features
di: Paulo, Gonçalo, et al.
Pubblicazione: (2025)
di: Paulo, Gonçalo, et al.
Pubblicazione: (2025)
Does Transformer Interpretability Transfer to RNNs?
di: Paulo, Gonçalo, et al.
Pubblicazione: (2024)
di: Paulo, Gonçalo, et al.
Pubblicazione: (2024)
Estimating the Probability of Sampling a Trained Neural Network at Random
di: Scherlis, Adam, et al.
Pubblicazione: (2025)
di: Scherlis, Adam, et al.
Pubblicazione: (2025)
Evaluating SAE interpretability without explanations
di: Paulo, Gonçalo, et al.
Pubblicazione: (2025)
di: Paulo, Gonçalo, et al.
Pubblicazione: (2025)
Converting MLPs into Polynomials in Closed Form
di: Belrose, Nora, et al.
Pubblicazione: (2025)
di: Belrose, Nora, et al.
Pubblicazione: (2025)
Balancing Label Quantity and Quality for Scalable Elicitation
di: Mallen, Alex, et al.
Pubblicazione: (2024)
di: Mallen, Alex, et al.
Pubblicazione: (2024)
Understanding Gradient Descent through the Training Jacobian
di: Belrose, Nora, et al.
Pubblicazione: (2024)
di: Belrose, Nora, et al.
Pubblicazione: (2024)
Automatically Interpreting Millions of Features in Large Language Models
di: Paulo, Gonçalo, et al.
Pubblicazione: (2024)
di: Paulo, Gonçalo, et al.
Pubblicazione: (2024)
Partially Rewriting a Transformer in Natural Language
di: Paulo, Gonçalo, et al.
Pubblicazione: (2025)
di: Paulo, Gonçalo, et al.
Pubblicazione: (2025)
Examining Two Hop Reasoning Through Information Content Scaling
di: Johnston, David, et al.
Pubblicazione: (2025)
di: Johnston, David, et al.
Pubblicazione: (2025)
Interpreting Large Text-to-Image Diffusion Models with Dictionary Learning
di: Shabalin, Stepan, et al.
Pubblicazione: (2025)
di: Shabalin, Stepan, et al.
Pubblicazione: (2025)
Refusal in LLMs is an Affine Function
di: Marshall, Thomas, et al.
Pubblicazione: (2024)
di: Marshall, Thomas, et al.
Pubblicazione: (2024)
Mechanistic Anomaly Detection for "Quirky" Language Models
di: Johnston, David O., et al.
Pubblicazione: (2025)
di: Johnston, David O., et al.
Pubblicazione: (2025)
Scaling sparse feature circuit finding for in-context learning
di: Kharlapenko, Dmitrii, et al.
Pubblicazione: (2025)
di: Kharlapenko, Dmitrii, et al.
Pubblicazione: (2025)
Eliciting Latent Knowledge from Quirky Language Models
di: Mallen, Alex, et al.
Pubblicazione: (2023)
di: Mallen, Alex, et al.
Pubblicazione: (2023)
Understanding Addition in Transformers
di: Quirke, Philip, et al.
Pubblicazione: (2023)
di: Quirke, Philip, et al.
Pubblicazione: (2023)
TinySQL: A Progressive Text-to-SQL Dataset for Mechanistic Interpretability Research
di: Harrasse, Abir, et al.
Pubblicazione: (2025)
di: Harrasse, Abir, et al.
Pubblicazione: (2025)
Understanding Addition and Subtraction in Transformers
di: Quirke, Philip, et al.
Pubblicazione: (2024)
di: Quirke, Philip, et al.
Pubblicazione: (2024)
Mechanistic Interpretability of Code Correctness in LLMs via Sparse Autoencoders
di: Tahimic, Kriz, et al.
Pubblicazione: (2025)
di: Tahimic, Kriz, et al.
Pubblicazione: (2025)
Competition Report: Finding Universal Jailbreak Backdoors in Aligned LLMs
di: Rando, Javier, et al.
Pubblicazione: (2024)
di: Rando, Javier, et al.
Pubblicazione: (2024)
LEACE: Perfect linear concept erasure in closed form
di: Belrose, Nora, et al.
Pubblicazione: (2023)
di: Belrose, Nora, et al.
Pubblicazione: (2023)
Eliciting Latent Predictions from Transformers with the Tuned Lens
di: Belrose, Nora, et al.
Pubblicazione: (2023)
di: Belrose, Nora, et al.
Pubblicazione: (2023)
Sparse Binary Representation Learning for Knowledge Tracing
di: Badran, Yahya, et al.
Pubblicazione: (2025)
di: Badran, Yahya, et al.
Pubblicazione: (2025)
Seeking Interpretability and Explainability in Binary Activated Neural Networks
di: Leblanc, Benjamin, et al.
Pubblicazione: (2022)
di: Leblanc, Benjamin, et al.
Pubblicazione: (2022)
Interpretable Reward Model via Sparse Autoencoder
di: Zhang, Shuyi, et al.
Pubblicazione: (2025)
di: Zhang, Shuyi, et al.
Pubblicazione: (2025)
Interpreting Attention Layer Outputs with Sparse Autoencoders
di: Kissane, Connor, et al.
Pubblicazione: (2024)
di: Kissane, Connor, et al.
Pubblicazione: (2024)
Mechanistic Interpretability of Binary and Ternary Transformers
di: Li, Jason
Pubblicazione: (2024)
di: Li, Jason
Pubblicazione: (2024)
Less Discriminatory Alternative and Interpretable XGBoost Framework for Binary Classification
di: Pangia, Andrew, et al.
Pubblicazione: (2024)
di: Pangia, Andrew, et al.
Pubblicazione: (2024)
Route Sparse Autoencoder to Interpret Large Language Models
di: Shi, Wei, et al.
Pubblicazione: (2025)
di: Shi, Wei, et al.
Pubblicazione: (2025)
Step-Level Sparse Autoencoder for Reasoning Process Interpretation
di: Yang, Xuan, et al.
Pubblicazione: (2026)
di: Yang, Xuan, et al.
Pubblicazione: (2026)
Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control
di: Makelov, Aleksandar, et al.
Pubblicazione: (2024)
di: Makelov, Aleksandar, et al.
Pubblicazione: (2024)
An Interpretable and Stable Framework for Sparse Principal Component Analysis
di: Hu, Ying, et al.
Pubblicazione: (2026)
di: Hu, Ying, et al.
Pubblicazione: (2026)
Sparse Coding Representation of 2-way Data
di: Ma, Boya, et al.
Pubblicazione: (2025)
di: Ma, Boya, et al.
Pubblicazione: (2025)
HadamRNN: Binary and Sparse Ternary Orthogonal RNNs
di: Foucault, Armand, et al.
Pubblicazione: (2025)
di: Foucault, Armand, et al.
Pubblicazione: (2025)
Learning Sparse Codes with Entropy-Based ELBOs
di: Velychko, Dmytro, et al.
Pubblicazione: (2023)
di: Velychko, Dmytro, et al.
Pubblicazione: (2023)
Physics-Inspired Binary Neural Networks: Interpretable Compression with Theoretical Guarantees
di: Eamaz, Arian, et al.
Pubblicazione: (2025)
di: Eamaz, Arian, et al.
Pubblicazione: (2025)
Group Equivariance Meets Mechanistic Interpretability: Equivariant Sparse Autoencoders
di: Erdogan, Ege, et al.
Pubblicazione: (2025)
di: Erdogan, Ege, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Transcoders Beat Sparse Autoencoders for Interpretability
di: Paulo, Gonçalo, et al.
Pubblicazione: (2025) -
Slowing Learning by Erasing Simple Features
di: Quirke, Lucia, et al.
Pubblicazione: (2025) -
Neural Networks Learn Statistics of Increasing Complexity
di: Belrose, Nora, et al.
Pubblicazione: (2024) -
Sparse Autoencoders Trained on the Same Data Learn Different Features
di: Paulo, Gonçalo, et al.
Pubblicazione: (2025) -
Does Transformer Interpretability Transfer to RNNs?
di: Paulo, Gonçalo, et al.
Pubblicazione: (2024)