Distribution-Aware Feature Selection for SAEs
Fuente:
arXiv
Saved in:
| Main Authors: | Oozeer, Narmeen, Prakash, Nirmalendu, Lan, Michael, Rigg, Alice, Abdullah, Amirali |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bilinear Convolution Decomposition for Causal RL Interpretability
by: Oozeer, Narmeen, et al.
Published: (2024)
by: Oozeer, Narmeen, et al.
Published: (2024)
DreamReader: An Interpretability Toolkit for Text-to-Image Models
by: Prakash, Nirmalendu, et al.
Published: (2026)
by: Prakash, Nirmalendu, et al.
Published: (2026)
Activation Space Interventions Can Be Transferred Between Large Language Models
by: Oozeer, Narmeen, et al.
Published: (2025)
by: Oozeer, Narmeen, et al.
Published: (2025)
Beyond Linear Steering: Unified Multi-Attribute Control for Language Models
by: Oozeer, Narmeen, et al.
Published: (2025)
by: Oozeer, Narmeen, et al.
Published: (2025)
Understanding and Mitigating Dataset Corruption in LLM Steering
by: Anderson, Cullen, et al.
Published: (2026)
by: Anderson, Cullen, et al.
Published: (2026)
Spectral Superposition: A Theory of Feature Geometry
by: Ivanov, Georgi, et al.
Published: (2026)
by: Ivanov, Georgi, et al.
Published: (2026)
SAEs Are Good for Steering -- If You Select the Right Features
by: Arad, Dana, et al.
Published: (2025)
by: Arad, Dana, et al.
Published: (2025)
Converting MLPs into Polynomials in Closed Form
by: Belrose, Nora, et al.
Published: (2025)
by: Belrose, Nora, et al.
Published: (2025)
Weight-based Decomposition: A Case for Bilinear MLPs
by: Pearce, Michael T., et al.
Published: (2024)
by: Pearce, Michael T., et al.
Published: (2024)
Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators
by: Roytburg, Dani, et al.
Published: (2025)
by: Roytburg, Dani, et al.
Published: (2025)
Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations
by: Roytburg, Dani, et al.
Published: (2026)
by: Roytburg, Dani, et al.
Published: (2026)
Features Emerge as Discrete States: The First Application of SAEs to 3D Representations
by: Miao, Albert, et al.
Published: (2025)
by: Miao, Albert, et al.
Published: (2025)
Tokenized SAEs: Disentangling SAE Reconstructions
by: Dooms, Thomas, et al.
Published: (2025)
by: Dooms, Thomas, et al.
Published: (2025)
The Rate-Distortion-Polysemanticity Tradeoff in SAEs
by: Mencattini, Tommaso, et al.
Published: (2026)
by: Mencattini, Tommaso, et al.
Published: (2026)
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs
by: Song, Xiangchen, et al.
Published: (2025)
by: Song, Xiangchen, et al.
Published: (2025)
Analyzing (In)Abilities of SAEs via Formal Languages
by: Menon, Abhinav, et al.
Published: (2024)
by: Menon, Abhinav, et al.
Published: (2024)
Bilinear MLPs enable weight-based mechanistic interpretability
by: Pearce, Michael T., et al.
Published: (2024)
by: Pearce, Michael T., et al.
Published: (2024)
Residual Stream Analysis with Multi-Layer SAEs
by: Lawson, Tim, et al.
Published: (2024)
by: Lawson, Tim, et al.
Published: (2024)
Approximating Human Preferences Using a Multi-Judge Learned System
by: Sprejer, Eitán, et al.
Published: (2025)
by: Sprejer, Eitán, et al.
Published: (2025)
Sanity Checks for Sparse Autoencoders: Do SAEs Beat Random Baselines?
by: Korznikov, Anton, et al.
Published: (2026)
by: Korznikov, Anton, et al.
Published: (2026)
Ablating Archetypes: The Stability of Archetypal SAEs is an Artifact of Initialization and Metric Design
by: Brzozowski, Michał, et al.
Published: (2026)
by: Brzozowski, Michał, et al.
Published: (2026)
Resa: Transparent Reasoning Models via SAEs
by: Wang, Shangshang, et al.
Published: (2025)
by: Wang, Shangshang, et al.
Published: (2025)
Interpretability as Compression: Reconsidering SAE Explanations of Neural Activations with MDL-SAEs
by: Ayonrinde, Kola, et al.
Published: (2024)
by: Ayonrinde, Kola, et al.
Published: (2024)
Structure-Aware Masking for Protein Representation Learning
by: Walton, Thomas, et al.
Published: (2026)
by: Walton, Thomas, et al.
Published: (2026)
Can SAEs reveal and mitigate racial biases of LLMs in healthcare?
by: Ahsan, Hiba, et al.
Published: (2025)
by: Ahsan, Hiba, et al.
Published: (2025)
Investigating Sensitive Directions in GPT-2: An Improved Baseline and Comparative Analysis of SAEs
by: Lee, Daniel J., et al.
Published: (2024)
by: Lee, Daniel J., et al.
Published: (2024)
Teach Old SAEs New Domain Tricks with Boosting
by: Koriagin, Nikita, et al.
Published: (2025)
by: Koriagin, Nikita, et al.
Published: (2025)
Distributionally Robust Feature Selection
by: Swaroop, Maitreyi, et al.
Published: (2025)
by: Swaroop, Maitreyi, et al.
Published: (2025)
Detecting and Characterizing Planning in Language Models
by: Nainani, Jatin, et al.
Published: (2025)
by: Nainani, Jatin, et al.
Published: (2025)
Geometry-Aware CLIP Retrieval via Local Cross-Modal Alignment and Steering
by: Prakash, Nirmalendu, et al.
Published: (2026)
by: Prakash, Nirmalendu, et al.
Published: (2026)
A Context-Aware Temporal Modeling through Unified Multi-Scale Temporal Encoding and Hierarchical Sequence Learning for Single-Channel EEG Sleep Staging
by: Vakili, Amirali, et al.
Published: (2025)
by: Vakili, Amirali, et al.
Published: (2025)
Mechanistic Interpretability with SAEs: Probing Religion, Violence, and Geography in Large Language Models
by: Simbeck, Katharina, et al.
Published: (2025)
by: Simbeck, Katharina, et al.
Published: (2025)
Causally-Aware Unsupervised Feature Selection Learning
by: Shen, Zongxin, et al.
Published: (2024)
by: Shen, Zongxin, et al.
Published: (2024)
A Multicollinearity-Aware Signal-Processing Framework for Cross-$β$ Identification via X-ray Scattering of Alzheimer's Tissue
by: Bashit, Abdullah Al, et al.
Published: (2025)
by: Bashit, Abdullah Al, et al.
Published: (2025)
Sycophancy as compositions of Atomic Psychometric Traits
by: Jain, Shreyans, et al.
Published: (2025)
by: Jain, Shreyans, et al.
Published: (2025)
Minimax and Communication-Efficient Distributed Best Subset Selection with Oracle Property
by: Lan, Jingguo, et al.
Published: (2024)
by: Lan, Jingguo, et al.
Published: (2024)
Safe Distributionally Robust Feature Selection under Covariate Shift
by: Hanada, Hiroyuki, et al.
Published: (2026)
by: Hanada, Hiroyuki, et al.
Published: (2026)
Effective Feature Selection for Predicting Spreading Factor with ML in Large LoRaWAN-based Mobile IoT Networks
by: Prakash, Aman, et al.
Published: (2025)
by: Prakash, Aman, et al.
Published: (2025)
Scalable Expressiveness through Preprocessed Graph Perturbations
by: Saber, Danial, et al.
Published: (2024)
by: Saber, Danial, et al.
Published: (2024)
RADE: Random Add-Drop Edge as a Regularizer
by: Saber, Danial, et al.
Published: (2026)
by: Saber, Danial, et al.
Published: (2026)
Similar Items
-
Bilinear Convolution Decomposition for Causal RL Interpretability
by: Oozeer, Narmeen, et al.
Published: (2024) -
DreamReader: An Interpretability Toolkit for Text-to-Image Models
by: Prakash, Nirmalendu, et al.
Published: (2026) -
Activation Space Interventions Can Be Transferred Between Large Language Models
by: Oozeer, Narmeen, et al.
Published: (2025) -
Beyond Linear Steering: Unified Multi-Attribute Control for Language Models
by: Oozeer, Narmeen, et al.
Published: (2025) -
Understanding and Mitigating Dataset Corruption in LLM Steering
by: Anderson, Cullen, et al.
Published: (2026)