Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lieberum, Tom, Rajamanoharan, Senthooran, Conmy, Arthur, Smith, Lewis, Sonnerat, Nicolas, Varma, Vikrant, Kramár, János, Dragan, Anca, Shah, Rohin, Nanda, Neel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders
von: Rajamanoharan, Senthooran, et al.
Veröffentlicht: (2024)
von: Rajamanoharan, Senthooran, et al.
Veröffentlicht: (2024)
Improving Dictionary Learning with Gated Sparse Autoencoders
von: Rajamanoharan, Senthooran, et al.
Veröffentlicht: (2024)
von: Rajamanoharan, Senthooran, et al.
Veröffentlicht: (2024)
AtP*: An efficient and scalable method for localizing LLM behaviour to components
von: Kramár, János, et al.
Veröffentlicht: (2024)
von: Kramár, János, et al.
Veröffentlicht: (2024)
Subliminal Learning Is Steering Vector Distillation
von: Blank, Camila, et al.
Veröffentlicht: (2026)
von: Blank, Camila, et al.
Veröffentlicht: (2026)
Building Production-Ready Probes For Gemini
von: Kramár, János, et al.
Veröffentlicht: (2026)
von: Kramár, János, et al.
Veröffentlicht: (2026)
Are Sparse Autoencoders Useful? A Case Study in Sparse Probing
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2025)
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2025)
Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
von: Arcuschin, Iván, et al.
Veröffentlicht: (2025)
von: Arcuschin, Iván, et al.
Veröffentlicht: (2025)
How Well Do Models Follow Their Constitutions?
von: Jakkli, Arya, et al.
Veröffentlicht: (2026)
von: Jakkli, Arya, et al.
Veröffentlicht: (2026)
Eliciting Secret Knowledge from Language Models
von: Cywiński, Bartosz, et al.
Veröffentlicht: (2025)
von: Cywiński, Bartosz, et al.
Veröffentlicht: (2025)
Towards eliciting latent knowledge from LLMs with mechanistic interpretability
von: Cywiński, Bartosz, et al.
Veröffentlicht: (2025)
von: Cywiński, Bartosz, et al.
Veröffentlicht: (2025)
Emergent Misalignment is Easy, Narrow Misalignment is Hard
von: Soligo, Anna, et al.
Veröffentlicht: (2026)
von: Soligo, Anna, et al.
Veröffentlicht: (2026)
Do I Know This Entity? Knowledge Awareness and Hallucinations in Language Models
von: Ferrando, Javier, et al.
Veröffentlicht: (2024)
von: Ferrando, Javier, et al.
Veröffentlicht: (2024)
Convergent Linear Representations of Emergent Misalignment
von: Soligo, Anna, et al.
Veröffentlicht: (2025)
von: Soligo, Anna, et al.
Veröffentlicht: (2025)
Interpreting Attention Layer Outputs with Sparse Autoencoders
von: Kissane, Connor, et al.
Veröffentlicht: (2024)
von: Kissane, Connor, et al.
Veröffentlicht: (2024)
Thought Branches: Interpreting LLM Reasoning Requires Resampling
von: Macar, Uzay, et al.
Veröffentlicht: (2025)
von: Macar, Uzay, et al.
Veröffentlicht: (2025)
Model Organisms for Emergent Misalignment
von: Turner, Edward, et al.
Veröffentlicht: (2025)
von: Turner, Edward, et al.
Veröffentlicht: (2025)
Simple Mechanistic Explanations for Out-Of-Context Reasoning
von: Wang, Atticus, et al.
Veröffentlicht: (2025)
von: Wang, Atticus, et al.
Veröffentlicht: (2025)
Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning
von: Casademunt, Helena, et al.
Veröffentlicht: (2025)
von: Casademunt, Helena, et al.
Veröffentlicht: (2025)
CodeGemma: Open Code Models Based on Gemma
von: CodeGemma Team, et al.
Veröffentlicht: (2024)
von: CodeGemma Team, et al.
Veröffentlicht: (2024)
VaultGemma: A Differentially Private Gemma Model
von: Sinha, Amer, et al.
Veröffentlicht: (2025)
von: Sinha, Amer, et al.
Veröffentlicht: (2025)
ShieldGemma: Generative AI Content Moderation Based on Gemma
von: Zeng, Wenjun, et al.
Veröffentlicht: (2024)
von: Zeng, Wenjun, et al.
Veröffentlicht: (2024)
Improving Steering Vectors by Targeting Sparse Autoencoder Features
von: Chalnev, Sviatoslav, et al.
Veröffentlicht: (2024)
von: Chalnev, Sviatoslav, et al.
Veröffentlicht: (2024)
BatchTopK Sparse Autoencoders
von: Bussmann, Bart, et al.
Veröffentlicht: (2024)
von: Bussmann, Bart, et al.
Veröffentlicht: (2024)
An Approach to Technical AGI Safety and Security
von: Shah, Rohin, et al.
Veröffentlicht: (2025)
von: Shah, Rohin, et al.
Veröffentlicht: (2025)
Thought Anchors: Which LLM Reasoning Steps Matter?
von: Bogdan, Paul C., et al.
Veröffentlicht: (2025)
von: Bogdan, Paul C., et al.
Veröffentlicht: (2025)
Interpretable Embeddings with Sparse Autoencoders: A Data Analysis Toolkit
von: Jiang, Nick, et al.
Veröffentlicht: (2025)
von: Jiang, Nick, et al.
Veröffentlicht: (2025)
Amplitude Uncertainties Everywhere All at Once
von: Bahl, Henning, et al.
Veröffentlicht: (2025)
von: Bahl, Henning, et al.
Veröffentlicht: (2025)
Pruning Everything, Everywhere, All at Once
von: Nascimento, Gustavo Henrique do, et al.
Veröffentlicht: (2025)
von: Nascimento, Gustavo Henrique do, et al.
Veröffentlicht: (2025)
Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control
von: Makelov, Aleksandar, et al.
Veröffentlicht: (2024)
von: Makelov, Aleksandar, et al.
Veröffentlicht: (2024)
Every Centiloid, from Everywhere, All at Once
von: Ganna Blazhenets, et al.
Veröffentlicht: (2025)
von: Ganna Blazhenets, et al.
Veröffentlicht: (2025)
Every Centiloid, from Everywhere, All at Once
von: Ganna Blazhenets, et al.
Veröffentlicht: (2025)
von: Ganna Blazhenets, et al.
Veröffentlicht: (2025)
Understanding Reasoning in Thinking Language Models via Steering Vectors
von: Venhoff, Constantin, et al.
Veröffentlicht: (2025)
von: Venhoff, Constantin, et al.
Veröffentlicht: (2025)
Scaling sparse feature circuit finding for in-context learning
von: Kharlapenko, Dmitrii, et al.
Veröffentlicht: (2025)
von: Kharlapenko, Dmitrii, et al.
Veröffentlicht: (2025)
Base Models Know How to Reason, Thinking Models Learn When
von: Venhoff, Constantin, et al.
Veröffentlicht: (2025)
von: Venhoff, Constantin, et al.
Veröffentlicht: (2025)
Gemma-2-2b
von: Raj, Suyash
Veröffentlicht: (2025)
von: Raj, Suyash
Veröffentlicht: (2025)
TranslateGemma Technical Report
von: Finkelstein, Mara, et al.
Veröffentlicht: (2026)
von: Finkelstein, Mara, et al.
Veröffentlicht: (2026)
MedGemma Technical Report
von: Sellergren, Andrew, et al.
Veröffentlicht: (2025)
von: Sellergren, Andrew, et al.
Veröffentlicht: (2025)
Gemma 3 Technical Report
von: Gemma Team, et al.
Veröffentlicht: (2025)
von: Gemma Team, et al.
Veröffentlicht: (2025)
Learning Multi-Level Features with Matryoshka Sparse Autoencoders
von: Bussmann, Bart, et al.
Veröffentlicht: (2025)
von: Bussmann, Bart, et al.
Veröffentlicht: (2025)
Evaluating Sparse Autoencoders on Targeted Concept Erasure Tasks
von: Karvonen, Adam, et al.
Veröffentlicht: (2024)
von: Karvonen, Adam, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders
von: Rajamanoharan, Senthooran, et al.
Veröffentlicht: (2024) -
Improving Dictionary Learning with Gated Sparse Autoencoders
von: Rajamanoharan, Senthooran, et al.
Veröffentlicht: (2024) -
AtP*: An efficient and scalable method for localizing LLM behaviour to components
von: Kramár, János, et al.
Veröffentlicht: (2024) -
Subliminal Learning Is Steering Vector Distillation
von: Blank, Camila, et al.
Veröffentlicht: (2026) -
Building Production-Ready Probes For Gemini
von: Kramár, János, et al.
Veröffentlicht: (2026)