A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chanin, David, Wilken-Smith, James, Dulka, Tomáš, Bhatnagar, Hardik, Golechha, Satvik, Bloom, Joseph |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Feature Hedging: Correlated Features Break Narrow Sparse Autoencoders
von: Chanin, David, et al.
Veröffentlicht: (2025)
von: Chanin, David, et al.
Veröffentlicht: (2025)
Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders
von: Chanin, David, et al.
Veröffentlicht: (2025)
von: Chanin, David, et al.
Veröffentlicht: (2025)
NICE: To Optimize In-Context Examples or Not?
von: Srivastava, Pragya, et al.
Veröffentlicht: (2024)
von: Srivastava, Pragya, et al.
Veröffentlicht: (2024)
Are Sparse Autoencoder Benchmarks Reliable?
von: Chanin, David
Veröffentlicht: (2026)
von: Chanin, David
Veröffentlicht: (2026)
Progress Measures for Grokking on Real-world Tasks
von: Golechha, Satvik
Veröffentlicht: (2024)
von: Golechha, Satvik
Veröffentlicht: (2024)
ABBEL: LLM Agents Acting through Belief Bottlenecks Expressed in Language
von: Lidayan, Aly, et al.
Veröffentlicht: (2025)
von: Lidayan, Aly, et al.
Veröffentlicht: (2025)
Challenges in Mechanistically Interpreting Model Representations
von: Golechha, Satvik, et al.
Veröffentlicht: (2024)
von: Golechha, Satvik, et al.
Veröffentlicht: (2024)
Building Better Deception Probes Using Targeted Instruction Pairs
von: Natarajan, Vikram, et al.
Veröffentlicht: (2026)
von: Natarajan, Vikram, et al.
Veröffentlicht: (2026)
Sparse Autoencoder Features for Classifications and Transferability
von: Gallifant, Jack, et al.
Veröffentlicht: (2025)
von: Gallifant, Jack, et al.
Veröffentlicht: (2025)
Identifying Linear Relational Concepts in Large Language Models
von: Chanin, David, et al.
Veröffentlicht: (2023)
von: Chanin, David, et al.
Veröffentlicht: (2023)
Among Us: A Sandbox for Measuring and Detecting Agentic Deception
von: Golechha, Satvik, et al.
Veröffentlicht: (2025)
von: Golechha, Satvik, et al.
Veröffentlicht: (2025)
Improving Steering Vectors by Targeting Sparse Autoencoder Features
von: Chalnev, Sviatoslav, et al.
Veröffentlicht: (2024)
von: Chalnev, Sviatoslav, et al.
Veröffentlicht: (2024)
SynthSAEBench: Evaluating Sparse Autoencoders on Scalable Realistic Synthetic Data
von: Chanin, David, et al.
Veröffentlicht: (2026)
von: Chanin, David, et al.
Veröffentlicht: (2026)
SparseRM: A Lightweight Preference Modeling with Sparse Autoencoder
von: Liu, Dengcan, et al.
Veröffentlicht: (2025)
von: Liu, Dengcan, et al.
Veröffentlicht: (2025)
Constrain Alignment with Sparse Autoencoders
von: Yin, Qingyu, et al.
Veröffentlicht: (2024)
von: Yin, Qingyu, et al.
Veröffentlicht: (2024)
AbsTopK: Rethinking Sparse Autoencoders For Bidirectional Features
von: Zhu, Xudong, et al.
Veröffentlicht: (2025)
von: Zhu, Xudong, et al.
Veröffentlicht: (2025)
Mechanistic Knobs in LLMs: Retrieving and Steering High-Order Semantic Features via Sparse Autoencoders
von: Zhang, Ruikang, et al.
Veröffentlicht: (2026)
von: Zhang, Ruikang, et al.
Veröffentlicht: (2026)
Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders
von: Lan, Michael, et al.
Veröffentlicht: (2024)
von: Lan, Michael, et al.
Veröffentlicht: (2024)
Sparse Autoencoders for Hypothesis Generation
von: Movva, Rajiv, et al.
Veröffentlicht: (2025)
von: Movva, Rajiv, et al.
Veröffentlicht: (2025)
How Do LLMs Encode Scientific Quality? An Empirical Study Using Monosemantic Features from Sparse Autoencoders
von: McCoubrey, Michael, et al.
Veröffentlicht: (2026)
von: McCoubrey, Michael, et al.
Veröffentlicht: (2026)
Training Neural Networks for Modularity aids Interpretability
von: Golechha, Satvik, et al.
Veröffentlicht: (2024)
von: Golechha, Satvik, et al.
Veröffentlicht: (2024)
SAFER: Probing Safety in Reward Models with Sparse Autoencoder
von: Shi, Wei, et al.
Veröffentlicht: (2025)
von: Shi, Wei, et al.
Veröffentlicht: (2025)
Toward Faithful Retrieval-Augmented Generation with Sparse Autoencoders
von: Xiong, Guangzhi, et al.
Veröffentlicht: (2025)
von: Xiong, Guangzhi, et al.
Veröffentlicht: (2025)
Time-Aware Feature Selection: Adaptive Temporal Masking for Stable Sparse Autoencoder Training
von: Li, T. Ed, et al.
Veröffentlicht: (2025)
von: Li, T. Ed, et al.
Veröffentlicht: (2025)
Decoding Dark Matter: Specialized Sparse Autoencoders for Interpreting Rare Concepts in Foundation Models
von: Muhamed, Aashiq, et al.
Veröffentlicht: (2024)
von: Muhamed, Aashiq, et al.
Veröffentlicht: (2024)
Controllable LLM Reasoning via Sparse Autoencoder-Based Steering
von: Fang, Yi, et al.
Veröffentlicht: (2026)
von: Fang, Yi, et al.
Veröffentlicht: (2026)
Disentangling concept semantics via multilingual averaging in Sparse Autoencoders
von: O'Reilly, Cliff, et al.
Veröffentlicht: (2025)
von: O'Reilly, Cliff, et al.
Veröffentlicht: (2025)
Uncovering Cross-Linguistic Disparities in LLMs using Sparse Autoencoders
von: Xuan, Richmond Sin Jing, et al.
Veröffentlicht: (2025)
von: Xuan, Richmond Sin Jing, et al.
Veröffentlicht: (2025)
Dissecting Chronos: Sparse Autoencoders Reveal Causal Feature Hierarchies in Time Series Foundation Models
von: Mishra, Anurag
Veröffentlicht: (2026)
von: Mishra, Anurag
Veröffentlicht: (2026)
FaithfulSAE: Towards Capturing Faithful Features with Sparse Autoencoders without External Dataset Dependencies
von: Cho, Seonglae, et al.
Veröffentlicht: (2025)
von: Cho, Seonglae, et al.
Veröffentlicht: (2025)
Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering
von: Zhao, Haiyan, et al.
Veröffentlicht: (2025)
von: Zhao, Haiyan, et al.
Veröffentlicht: (2025)
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation
von: Shu, Huizhen, et al.
Veröffentlicht: (2025)
von: Shu, Huizhen, et al.
Veröffentlicht: (2025)
Sparse Autoencoder Decomposition of Clinical Sequence Model Representations: Feature Complexity, Task Specialisation, and Mortality Prediction
von: Sainsbury, Chris, et al.
Veröffentlicht: (2026)
von: Sainsbury, Chris, et al.
Veröffentlicht: (2026)
Enabling Precise Topic Alignment in Large Language Models Via Sparse Autoencoders
von: Joshi, Ananya, et al.
Veröffentlicht: (2025)
von: Joshi, Ananya, et al.
Veröffentlicht: (2025)
Incorporating Hierarchical Semantics in Sparse Autoencoder Architectures
von: Muchane, Mark, et al.
Veröffentlicht: (2025)
von: Muchane, Mark, et al.
Veröffentlicht: (2025)
Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
von: Lieberum, Tom, et al.
Veröffentlicht: (2024)
von: Lieberum, Tom, et al.
Veröffentlicht: (2024)
SMoA: Improving Multi-agent Large Language Models with Sparse Mixture-of-Agents
von: Li, Dawei, et al.
Veröffentlicht: (2024)
von: Li, Dawei, et al.
Veröffentlicht: (2024)
Lost in Translation? A Comparative Study on the Cross-Lingual Transfer of Composite Harms
von: Shukla, Vaibhav, et al.
Veröffentlicht: (2026)
von: Shukla, Vaibhav, et al.
Veröffentlicht: (2026)
Jacobian Sparse Autoencoders: Sparsify Computations, Not Just Activations
von: Farnik, Lucy, et al.
Veröffentlicht: (2025)
von: Farnik, Lucy, et al.
Veröffentlicht: (2025)
Evaluating Adversarial Robustness of Concept Representations in Sparse Autoencoders
von: Li, Aaron J., et al.
Veröffentlicht: (2025)
von: Li, Aaron J., et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Feature Hedging: Correlated Features Break Narrow Sparse Autoencoders
von: Chanin, David, et al.
Veröffentlicht: (2025) -
Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders
von: Chanin, David, et al.
Veröffentlicht: (2025) -
NICE: To Optimize In-Context Examples or Not?
von: Srivastava, Pragya, et al.
Veröffentlicht: (2024) -
Are Sparse Autoencoder Benchmarks Reliable?
von: Chanin, David
Veröffentlicht: (2026) -
Progress Measures for Grokking on Real-world Tasks
von: Golechha, Satvik
Veröffentlicht: (2024)