Kronecker Factorization Improves Efficiency and Interpretability of Sparse Autoencoders
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kurochkin, Vadim, Aksenov, Yaroslav, Laptev, Daniil, Gavrilov, Daniil, Balagansky, Nikita |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Analyze Feature Flow to Enhance Interpretation and Steering in Language Models
von: Laptev, Daniil, et al.
Veröffentlicht: (2025)
von: Laptev, Daniil, et al.
Veröffentlicht: (2025)
Train One Sparse Autoencoder Across Multiple Sparsity Budgets to Preserve Interpretability and Accuracy
von: Balagansky, Nikita, et al.
Veröffentlicht: (2025)
von: Balagansky, Nikita, et al.
Veröffentlicht: (2025)
Teach Old SAEs New Domain Tricks with Boosting
von: Koriagin, Nikita, et al.
Veröffentlicht: (2025)
von: Koriagin, Nikita, et al.
Veröffentlicht: (2025)
You Do Not Fully Utilize Transformer's Representation Capacity
von: Gerasimov, Gleb, et al.
Veröffentlicht: (2025)
von: Gerasimov, Gleb, et al.
Veröffentlicht: (2025)
Small Vectors, Big Effects: A Mechanistic Study of RL-Induced Reasoning via Steering Vectors
von: Sinii, Viacheslav, et al.
Veröffentlicht: (2025)
von: Sinii, Viacheslav, et al.
Veröffentlicht: (2025)
Learn Your Reference Model for Real Good Alignment
von: Gorbatovski, Alexey, et al.
Veröffentlicht: (2024)
von: Gorbatovski, Alexey, et al.
Veröffentlicht: (2024)
Linear Transformers with Learnable Kernel Functions are Better In-Context Models
von: Aksenov, Yaroslav, et al.
Veröffentlicht: (2024)
von: Aksenov, Yaroslav, et al.
Veröffentlicht: (2024)
Diffusion Language Models Generation Can Be Halted Early
von: Vaina, Sofia Maria Lo Cicero, et al.
Veröffentlicht: (2023)
von: Vaina, Sofia Maria Lo Cicero, et al.
Veröffentlicht: (2023)
Mechanistic Permutability: Match Features Across Layers
von: Balagansky, Nikita, et al.
Veröffentlicht: (2024)
von: Balagansky, Nikita, et al.
Veröffentlicht: (2024)
Next Embedding Prediction Makes World Models Stronger
von: Bredis, George, et al.
Veröffentlicht: (2026)
von: Bredis, George, et al.
Veröffentlicht: (2026)
Trust-Region Behavior Blending for On-Policy Distillation
von: Plyusov, Daniil, et al.
Veröffentlicht: (2026)
von: Plyusov, Daniil, et al.
Veröffentlicht: (2026)
Steering LLM Reasoning Through Bias-Only Adaptation
von: Sinii, Viacheslav, et al.
Veröffentlicht: (2025)
von: Sinii, Viacheslav, et al.
Veröffentlicht: (2025)
Guided Star-Shaped Masked Diffusion
von: Meshchaninov, Viacheslav, et al.
Veröffentlicht: (2025)
von: Meshchaninov, Viacheslav, et al.
Veröffentlicht: (2025)
Interpretable Company Similarity with Sparse Autoencoders
von: Molinari, Marco, et al.
Veröffentlicht: (2024)
von: Molinari, Marco, et al.
Veröffentlicht: (2024)
SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability
von: Karvonen, Adam, et al.
Veröffentlicht: (2025)
von: Karvonen, Adam, et al.
Veröffentlicht: (2025)
PrivacyScalpel: Enhancing LLM Privacy via Interpretable Feature Intervention with Sparse Autoencoders
von: Frikha, Ahmed, et al.
Veröffentlicht: (2025)
von: Frikha, Ahmed, et al.
Veröffentlicht: (2025)
Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability
von: Bhalla, Usha, et al.
Veröffentlicht: (2025)
von: Bhalla, Usha, et al.
Veröffentlicht: (2025)
Improving Steering Vectors by Targeting Sparse Autoencoder Features
von: Chalnev, Sviatoslav, et al.
Veröffentlicht: (2024)
von: Chalnev, Sviatoslav, et al.
Veröffentlicht: (2024)
A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models
von: Shu, Dong, et al.
Veröffentlicht: (2025)
von: Shu, Dong, et al.
Veröffentlicht: (2025)
DLM-Scope: Mechanistic Interpretability of Diffusion Language Models via Sparse Autoencoders
von: Wang, Xu, et al.
Veröffentlicht: (2026)
von: Wang, Xu, et al.
Veröffentlicht: (2026)
Test Code Generation for Telecom Software Systems using Two-Stage Generative Model
von: Nabeel, Mohamad, et al.
Veröffentlicht: (2024)
von: Nabeel, Mohamad, et al.
Veröffentlicht: (2024)
Training Superior Sparse Autoencoders for Instruct Models
von: Li, Jiaming, et al.
Veröffentlicht: (2025)
von: Li, Jiaming, et al.
Veröffentlicht: (2025)
AlignSAE: Concept-Aligned Sparse Autoencoders
von: Yang, Minglai, et al.
Veröffentlicht: (2025)
von: Yang, Minglai, et al.
Veröffentlicht: (2025)
SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models
von: He, Zirui, et al.
Veröffentlicht: (2025)
von: He, Zirui, et al.
Veröffentlicht: (2025)
Decoding Dark Matter: Specialized Sparse Autoencoders for Interpreting Rare Concepts in Foundation Models
von: Muhamed, Aashiq, et al.
Veröffentlicht: (2024)
von: Muhamed, Aashiq, et al.
Veröffentlicht: (2024)
When an LLM is apprehensive about its answers -- and when its uncertainty is justified
von: Sychev, Petr, et al.
Veröffentlicht: (2025)
von: Sychev, Petr, et al.
Veröffentlicht: (2025)
Complexity-aware fine-tuning
von: Goncharov, Andrey, et al.
Veröffentlicht: (2025)
von: Goncharov, Andrey, et al.
Veröffentlicht: (2025)
Evaluating Sparse Autoencoders on Targeted Concept Erasure Tasks
von: Karvonen, Adam, et al.
Veröffentlicht: (2024)
von: Karvonen, Adam, et al.
Veröffentlicht: (2024)
MoRFI: Monotonic Sparse Autoencoder Feature Identification
von: Dimakopoulos, Dimitris, et al.
Veröffentlicht: (2026)
von: Dimakopoulos, Dimitris, et al.
Veröffentlicht: (2026)
Model Unlearning via Sparse Autoencoder Subspace Guided Projections
von: Wang, Xu, et al.
Veröffentlicht: (2025)
von: Wang, Xu, et al.
Veröffentlicht: (2025)
Visual Exploration of Feature Relationships in Sparse Autoencoders with Curated Concepts
von: Yan, Xinyuan, et al.
Veröffentlicht: (2025)
von: Yan, Xinyuan, et al.
Veröffentlicht: (2025)
DeMeVa at LeWiDi-2025: Modeling Perspectives with In-Context Learning and Label Distribution Learning
von: Ignatev, Daniil, et al.
Veröffentlicht: (2025)
von: Ignatev, Daniil, et al.
Veröffentlicht: (2025)
EEFSUVA: A New Mathematical Olympiad Benchmark
von: Khatibi, Nicole N, et al.
Veröffentlicht: (2025)
von: Khatibi, Nicole N, et al.
Veröffentlicht: (2025)
Sparse Autoencoder Features for Classifications and Transferability
von: Gallifant, Jack, et al.
Veröffentlicht: (2025)
von: Gallifant, Jack, et al.
Veröffentlicht: (2025)
Model Directions, Not Words: Mechanistic Topic Models Using Sparse Autoencoders
von: Zheng, Carolina, et al.
Veröffentlicht: (2025)
von: Zheng, Carolina, et al.
Veröffentlicht: (2025)
Sparse-Autoencoder-Guided Internal Representation Unlearning for Large Language Models
von: Yamashita, Tomoya, et al.
Veröffentlicht: (2025)
von: Yamashita, Tomoya, et al.
Veröffentlicht: (2025)
Sparse Autoencoders Enable Scalable and Reliable Circuit Identification in Language Models
von: O'Neill, Charles, et al.
Veröffentlicht: (2024)
von: O'Neill, Charles, et al.
Veröffentlicht: (2024)
Disentangling the Roles of Representation and Selection in Data Pruning
von: Du, Yupei, et al.
Veröffentlicht: (2025)
von: Du, Yupei, et al.
Veröffentlicht: (2025)
Incorporating Hierarchical Semantics in Sparse Autoencoder Architectures
von: Muchane, Mark, et al.
Veröffentlicht: (2025)
von: Muchane, Mark, et al.
Veröffentlicht: (2025)
Krony-PT: GPT2 compressed with Kronecker Products
von: Ayad, Mohamed Ayoub Ben, et al.
Veröffentlicht: (2024)
von: Ayad, Mohamed Ayoub Ben, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Analyze Feature Flow to Enhance Interpretation and Steering in Language Models
von: Laptev, Daniil, et al.
Veröffentlicht: (2025) -
Train One Sparse Autoencoder Across Multiple Sparsity Budgets to Preserve Interpretability and Accuracy
von: Balagansky, Nikita, et al.
Veröffentlicht: (2025) -
Teach Old SAEs New Domain Tricks with Boosting
von: Koriagin, Nikita, et al.
Veröffentlicht: (2025) -
You Do Not Fully Utilize Transformer's Representation Capacity
von: Gerasimov, Gleb, et al.
Veröffentlicht: (2025) -
Small Vectors, Big Effects: A Mechanistic Study of RL-Induced Reasoning via Steering Vectors
von: Sinii, Viacheslav, et al.
Veröffentlicht: (2025)