Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Rajamanoharan, Senthooran, Lieberum, Tom, Sonnerat, Nicolas, Conmy, Arthur, Varma, Vikrant, Kramár, János, Nanda, Neel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Improving Dictionary Learning with Gated Sparse Autoencoders
von: Rajamanoharan, Senthooran, et al.
Veröffentlicht: (2024)
von: Rajamanoharan, Senthooran, et al.
Veröffentlicht: (2024)
Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
von: Lieberum, Tom, et al.
Veröffentlicht: (2024)
von: Lieberum, Tom, et al.
Veröffentlicht: (2024)
Are Sparse Autoencoders Useful? A Case Study in Sparse Probing
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2025)
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2025)
AtP*: An efficient and scalable method for localizing LLM behaviour to components
von: Kramár, János, et al.
Veröffentlicht: (2024)
von: Kramár, János, et al.
Veröffentlicht: (2024)
Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
von: Arcuschin, Iván, et al.
Veröffentlicht: (2025)
von: Arcuschin, Iván, et al.
Veröffentlicht: (2025)
Subliminal Learning Is Steering Vector Distillation
von: Blank, Camila, et al.
Veröffentlicht: (2026)
von: Blank, Camila, et al.
Veröffentlicht: (2026)
Eliciting Secret Knowledge from Language Models
von: Cywiński, Bartosz, et al.
Veröffentlicht: (2025)
von: Cywiński, Bartosz, et al.
Veröffentlicht: (2025)
Towards eliciting latent knowledge from LLMs with mechanistic interpretability
von: Cywiński, Bartosz, et al.
Veröffentlicht: (2025)
von: Cywiński, Bartosz, et al.
Veröffentlicht: (2025)
Convergent Linear Representations of Emergent Misalignment
von: Soligo, Anna, et al.
Veröffentlicht: (2025)
von: Soligo, Anna, et al.
Veröffentlicht: (2025)
Do I Know This Entity? Knowledge Awareness and Hallucinations in Language Models
von: Ferrando, Javier, et al.
Veröffentlicht: (2024)
von: Ferrando, Javier, et al.
Veröffentlicht: (2024)
Interpreting Attention Layer Outputs with Sparse Autoencoders
von: Kissane, Connor, et al.
Veröffentlicht: (2024)
von: Kissane, Connor, et al.
Veröffentlicht: (2024)
Building Production-Ready Probes For Gemini
von: Kramár, János, et al.
Veröffentlicht: (2026)
von: Kramár, János, et al.
Veröffentlicht: (2026)
How Well Do Models Follow Their Constitutions?
von: Jakkli, Arya, et al.
Veröffentlicht: (2026)
von: Jakkli, Arya, et al.
Veröffentlicht: (2026)
Thought Branches: Interpreting LLM Reasoning Requires Resampling
von: Macar, Uzay, et al.
Veröffentlicht: (2025)
von: Macar, Uzay, et al.
Veröffentlicht: (2025)
Model Organisms for Emergent Misalignment
von: Turner, Edward, et al.
Veröffentlicht: (2025)
von: Turner, Edward, et al.
Veröffentlicht: (2025)
Sign-Aware Gated Sparse Autoencoders: Modeling Anticorrelated Features with Bi-Jump-ReLU Activations
von: Wieciech, Bartosz, et al.
Veröffentlicht: (2026)
von: Wieciech, Bartosz, et al.
Veröffentlicht: (2026)
Simple Mechanistic Explanations for Out-Of-Context Reasoning
von: Wang, Atticus, et al.
Veröffentlicht: (2025)
von: Wang, Atticus, et al.
Veröffentlicht: (2025)
Improving Steering Vectors by Targeting Sparse Autoencoder Features
von: Chalnev, Sviatoslav, et al.
Veröffentlicht: (2024)
von: Chalnev, Sviatoslav, et al.
Veröffentlicht: (2024)
Emergent Misalignment is Easy, Narrow Misalignment is Hard
von: Soligo, Anna, et al.
Veröffentlicht: (2026)
von: Soligo, Anna, et al.
Veröffentlicht: (2026)
Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning
von: Casademunt, Helena, et al.
Veröffentlicht: (2025)
von: Casademunt, Helena, et al.
Veröffentlicht: (2025)
BatchTopK Sparse Autoencoders
von: Bussmann, Bart, et al.
Veröffentlicht: (2024)
von: Bussmann, Bart, et al.
Veröffentlicht: (2024)
ReJump: A Tree-Jump Representation for Analyzing and Improving LLM Reasoning
von: Zeng, Yuchen, et al.
Veröffentlicht: (2025)
von: Zeng, Yuchen, et al.
Veröffentlicht: (2025)
Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control
von: Makelov, Aleksandar, et al.
Veröffentlicht: (2024)
von: Makelov, Aleksandar, et al.
Veröffentlicht: (2024)
Thought Anchors: Which LLM Reasoning Steps Matter?
von: Bogdan, Paul C., et al.
Veröffentlicht: (2025)
von: Bogdan, Paul C., et al.
Veröffentlicht: (2025)
Understanding Reasoning in Thinking Language Models via Steering Vectors
von: Venhoff, Constantin, et al.
Veröffentlicht: (2025)
von: Venhoff, Constantin, et al.
Veröffentlicht: (2025)
Base Models Know How to Reason, Thinking Models Learn When
von: Venhoff, Constantin, et al.
Veröffentlicht: (2025)
von: Venhoff, Constantin, et al.
Veröffentlicht: (2025)
Scaling sparse feature circuit finding for in-context learning
von: Kharlapenko, Dmitrii, et al.
Veröffentlicht: (2025)
von: Kharlapenko, Dmitrii, et al.
Veröffentlicht: (2025)
Learning Multi-Level Features with Matryoshka Sparse Autoencoders
von: Bussmann, Bart, et al.
Veröffentlicht: (2025)
von: Bussmann, Bart, et al.
Veröffentlicht: (2025)
Evaluating Sparse Autoencoders on Targeted Concept Erasure Tasks
von: Karvonen, Adam, et al.
Veröffentlicht: (2024)
von: Karvonen, Adam, et al.
Veröffentlicht: (2024)
Simple LLM Baselines are Competitive for Model Diffing
von: Kempf, Elias, et al.
Veröffentlicht: (2026)
von: Kempf, Elias, et al.
Veröffentlicht: (2026)
Interpretable Embeddings with Sparse Autoencoders: A Data Analysis Toolkit
von: Jiang, Nick, et al.
Veröffentlicht: (2025)
von: Jiang, Nick, et al.
Veröffentlicht: (2025)
SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability
von: Karvonen, Adam, et al.
Veröffentlicht: (2025)
von: Karvonen, Adam, et al.
Veröffentlicht: (2025)
ReSAE: Residualized Sparse Autoencoders for Multi-Layer Transformer Interventions
von: Poduval, Prathyush, et al.
Veröffentlicht: (2026)
von: Poduval, Prathyush, et al.
Veröffentlicht: (2026)
Sparse Autoencoders Do Not Find Canonical Units of Analysis
von: Leask, Patrick, et al.
Veröffentlicht: (2025)
von: Leask, Patrick, et al.
Veröffentlicht: (2025)
Modular Jump Gaussian Processes
von: Flowers, Anna R., et al.
Veröffentlicht: (2025)
von: Flowers, Anna R., et al.
Veröffentlicht: (2025)
An Approach to Technical AGI Safety and Security
von: Shah, Rohin, et al.
Veröffentlicht: (2025)
von: Shah, Rohin, et al.
Veröffentlicht: (2025)
Sparse Interaction Neighborhood Selection for Markov Random Fields via Reversible Jump and Pseudoposteriors
von: Freguglia, Victor, et al.
Veröffentlicht: (2022)
von: Freguglia, Victor, et al.
Veröffentlicht: (2022)
JumpLoRA: Sparse Adapters for Continual Learning in Large Language Models
von: Dragomir, Alexandra, et al.
Veröffentlicht: (2026)
von: Dragomir, Alexandra, et al.
Veröffentlicht: (2026)
Neural Jump ODEs as Generative Models
von: Crowell, Robert A., et al.
Veröffentlicht: (2025)
von: Crowell, Robert A., et al.
Veröffentlicht: (2025)
How to use and interpret activation patching
von: Heimersheim, Stefan, et al.
Veröffentlicht: (2024)
von: Heimersheim, Stefan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Improving Dictionary Learning with Gated Sparse Autoencoders
von: Rajamanoharan, Senthooran, et al.
Veröffentlicht: (2024) -
Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
von: Lieberum, Tom, et al.
Veröffentlicht: (2024) -
Are Sparse Autoencoders Useful? A Case Study in Sparse Probing
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2025) -
AtP*: An efficient and scalable method for localizing LLM behaviour to components
von: Kramár, János, et al.
Veröffentlicht: (2024) -
Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
von: Arcuschin, Iván, et al.
Veröffentlicht: (2025)