On the transferability of Sparse Autoencoders for interpreting compressed models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gupte, Suchit, Chhabra, Vishnu Kabir, Khalili, Mohammad Mahdi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Neuroplasticity and Corruption in Model Mechanisms: A Case Study Of Indirect Object Identification
von: Chhabra, Vishnu Kabir, et al.
Veröffentlicht: (2025)
von: Chhabra, Vishnu Kabir, et al.
Veröffentlicht: (2025)
Towards Understanding and Improving Refusal in Compressed Models via Mechanistic Interpretability
von: Chhabra, Vishnu Kabir, et al.
Veröffentlicht: (2025)
von: Chhabra, Vishnu Kabir, et al.
Veröffentlicht: (2025)
AbsTopK: Rethinking Sparse Autoencoders For Bidirectional Features
von: Zhu, Xudong, et al.
Veröffentlicht: (2025)
von: Zhu, Xudong, et al.
Veröffentlicht: (2025)
ECG Signal Denoising Using Multi-scale Patch Embedding and Transformers
von: Zhu, Ding, et al.
Veröffentlicht: (2024)
von: Zhu, Ding, et al.
Veröffentlicht: (2024)
Individual Fairness In Strategic Classification
von: Zuo, Zhiqun, et al.
Veröffentlicht: (2026)
von: Zuo, Zhiqun, et al.
Veröffentlicht: (2026)
Post-processing for Fair Regression via Explainable SVD
von: Zuo, Zhiqun, et al.
Veröffentlicht: (2025)
von: Zuo, Zhiqun, et al.
Veröffentlicht: (2025)
An Efficient Training Algorithm for Models with Block-wise Sparsity
von: Zhu, Ding, et al.
Veröffentlicht: (2025)
von: Zhu, Ding, et al.
Veröffentlicht: (2025)
CNN Autoencoders for Hierarchical Feature Extraction and Fusion in Multi-sensor Human Activity Recognition
von: Arabzadeh, Saeed, et al.
Veröffentlicht: (2025)
von: Arabzadeh, Saeed, et al.
Veröffentlicht: (2025)
Sparse Autoencoders, Again?
von: Lu, Yin, et al.
Veröffentlicht: (2025)
von: Lu, Yin, et al.
Veröffentlicht: (2025)
Revisiting the Effectiveness of LLM Pruning for Test-Time Scaling
von: Monjur, Ocean, et al.
Veröffentlicht: (2026)
von: Monjur, Ocean, et al.
Veröffentlicht: (2026)
Lookahead Counterfactual Fairness
von: Zuo, Zhiqun, et al.
Veröffentlicht: (2024)
von: Zuo, Zhiqun, et al.
Veröffentlicht: (2024)
Does higher interpretability imply better utility? A Pairwise Analysis on Sparse Autoencoders
von: Wang, Xu, et al.
Veröffentlicht: (2025)
von: Wang, Xu, et al.
Veröffentlicht: (2025)
Are Sparse Autoencoder Benchmarks Reliable?
von: Chanin, David
Veröffentlicht: (2026)
von: Chanin, David
Veröffentlicht: (2026)
From Emergence to Control: Probing and Modulating Self-Reflection in Language Models
von: Zhu, Xudong, et al.
Veröffentlicht: (2025)
von: Zhu, Xudong, et al.
Veröffentlicht: (2025)
BatchTopK Sparse Autoencoders
von: Bussmann, Bart, et al.
Veröffentlicht: (2024)
von: Bussmann, Bart, et al.
Veröffentlicht: (2024)
Improving Sparse Autoencoder with Dynamic Attention
von: Wang, Dongsheng, et al.
Veröffentlicht: (2026)
von: Wang, Dongsheng, et al.
Veröffentlicht: (2026)
Are Sparse Autoencoders Useful? A Case Study in Sparse Probing
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2025)
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2025)
Adaptive Sparse Allocation with Mutual Choice & Feature Choice Sparse Autoencoders
von: Ayonrinde, Kola
Veröffentlicht: (2024)
von: Ayonrinde, Kola
Veröffentlicht: (2024)
Less Diverse, Less Safe: The Indirect But Pervasive Risk of Test-Time Scaling in Large Language Models
von: Nahin, Shahriar Kabir, et al.
Veröffentlicht: (2025)
von: Nahin, Shahriar Kabir, et al.
Veröffentlicht: (2025)
Empirical Evaluation of Progressive Coding for Sparse Autoencoders
von: Peter, Hans, et al.
Veröffentlicht: (2025)
von: Peter, Hans, et al.
Veröffentlicht: (2025)
Data Whitening Improves Sparse Autoencoder Learning
von: Saraswatula, Ashwin, et al.
Veröffentlicht: (2025)
von: Saraswatula, Ashwin, et al.
Veröffentlicht: (2025)
Do Sparse Autoencoders Capture Concept Manifolds?
von: Bhalla, Usha, et al.
Veröffentlicht: (2026)
von: Bhalla, Usha, et al.
Veröffentlicht: (2026)
Improving Dictionary Learning with Gated Sparse Autoencoders
von: Rajamanoharan, Senthooran, et al.
Veröffentlicht: (2024)
von: Rajamanoharan, Senthooran, et al.
Veröffentlicht: (2024)
Learning Multi-Level Features with Matryoshka Sparse Autoencoders
von: Bussmann, Bart, et al.
Veröffentlicht: (2025)
von: Bussmann, Bart, et al.
Veröffentlicht: (2025)
Sparse Autoencoders Do Not Find Canonical Units of Analysis
von: Leask, Patrick, et al.
Veröffentlicht: (2025)
von: Leask, Patrick, et al.
Veröffentlicht: (2025)
Evaluating and Designing Sparse Autoencoders by Approximating Quasi-Orthogonality
von: Lee, Sewoong, et al.
Veröffentlicht: (2025)
von: Lee, Sewoong, et al.
Veröffentlicht: (2025)
Graph-Regularized Sparse Autoencoders for LLM Safety Steering
von: Yeon, Jehyeok, et al.
Veröffentlicht: (2025)
von: Yeon, Jehyeok, et al.
Veröffentlicht: (2025)
Improving Robustness In Sparse Autoencoders via Masked Regularization
von: Narayanaswamy, Vivek, et al.
Veröffentlicht: (2026)
von: Narayanaswamy, Vivek, et al.
Veröffentlicht: (2026)
Beyond Redundancy: Diverse and Specialized Multi-Expert Sparse Autoencoder
von: Xu, Zhen, et al.
Veröffentlicht: (2025)
von: Xu, Zhen, et al.
Veröffentlicht: (2025)
Projecting Assumptions: The Duality Between Sparse Autoencoders and Concept Geometry
von: Hindupur, Sai Sumedh R., et al.
Veröffentlicht: (2025)
von: Hindupur, Sai Sumedh R., et al.
Veröffentlicht: (2025)
Don't Forget It! Conditional Sparse Autoencoder Clamping Works for Unlearning
von: Khoriaty, Matthew, et al.
Veröffentlicht: (2025)
von: Khoriaty, Matthew, et al.
Veröffentlicht: (2025)
Interpretable Embeddings with Sparse Autoencoders: A Data Analysis Toolkit
von: Jiang, Nick, et al.
Veröffentlicht: (2025)
von: Jiang, Nick, et al.
Veröffentlicht: (2025)
Sparse Autoencoder Features for Classifications and Transferability
von: Gallifant, Jack, et al.
Veröffentlicht: (2025)
von: Gallifant, Jack, et al.
Veröffentlicht: (2025)
Learning Retrieval Models with Sparse Autoencoders
von: Formal, Thibault, et al.
Veröffentlicht: (2026)
von: Formal, Thibault, et al.
Veröffentlicht: (2026)
Finding Belief Geometries with Sparse Autoencoders
von: Levinson, Matthew
Veröffentlicht: (2026)
von: Levinson, Matthew
Veröffentlicht: (2026)
Superposition as Lossy Compression: Measure with Sparse Autoencoders and Connect to Adversarial Vulnerability
von: Bereska, Leonard, et al.
Veröffentlicht: (2025)
von: Bereska, Leonard, et al.
Veröffentlicht: (2025)
Enhancing LLM Steering through Sparse Autoencoder-Based Vector Refinement
von: Wang, Anyi, et al.
Veröffentlicht: (2025)
von: Wang, Anyi, et al.
Veröffentlicht: (2025)
SAeUron: Interpretable Concept Unlearning in Diffusion Models with Sparse Autoencoders
von: Cywiński, Bartosz, et al.
Veröffentlicht: (2025)
von: Cywiński, Bartosz, et al.
Veröffentlicht: (2025)
ReSAE: Residualized Sparse Autoencoders for Multi-Layer Transformer Interventions
von: Poduval, Prathyush, et al.
Veröffentlicht: (2026)
von: Poduval, Prathyush, et al.
Veröffentlicht: (2026)
Semantic Optimal Transport for Sparse Autoencoder Feature Matching and Circuit Compression
von: Cao, Tue M., et al.
Veröffentlicht: (2026)
von: Cao, Tue M., et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Neuroplasticity and Corruption in Model Mechanisms: A Case Study Of Indirect Object Identification
von: Chhabra, Vishnu Kabir, et al.
Veröffentlicht: (2025) -
Towards Understanding and Improving Refusal in Compressed Models via Mechanistic Interpretability
von: Chhabra, Vishnu Kabir, et al.
Veröffentlicht: (2025) -
AbsTopK: Rethinking Sparse Autoencoders For Bidirectional Features
von: Zhu, Xudong, et al.
Veröffentlicht: (2025) -
ECG Signal Denoising Using Multi-scale Patch Embedding and Transformers
von: Zhu, Ding, et al.
Veröffentlicht: (2024) -
Individual Fairness In Strategic Classification
von: Zuo, Zhiqun, et al.
Veröffentlicht: (2026)