Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models
Fuente:
arXiv
Saved in:
| Main Authors: | Karvonen, Adam, Wright, Benjamin, Rager, Can, Angell, Rico, Brinkmann, Jannik, Smith, Logan, Verdun, Claudio Mayrink, Bau, David, Marks, Samuel |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluating Sparse Autoencoders on Targeted Concept Erasure Tasks
by: Karvonen, Adam, et al.
Published: (2024)
by: Karvonen, Adam, et al.
Published: (2024)
Mitigating Adaptive Attacks against Reasoning Models with Activation Consistency Training
by: Shah, Avidan, et al.
Published: (2026)
by: Shah, Avidan, et al.
Published: (2026)
Jailbreak Transferability Emerges from Shared Representations
by: Angell, Rico, et al.
Published: (2025)
by: Angell, Rico, et al.
Published: (2025)
Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
by: Marks, Samuel, et al.
Published: (2024)
by: Marks, Samuel, et al.
Published: (2024)
Robustly Improving LLM Fairness in Realistic Settings via Interpretability
by: Karvonen, Adam, et al.
Published: (2025)
by: Karvonen, Adam, et al.
Published: (2025)
The Quest for the Right Mediator: Surveying Mechanistic Interpretability Through the Lens of Causal Mediation Analysis
by: Mueller, Aaron, et al.
Published: (2024)
by: Mueller, Aaron, et al.
Published: (2024)
Discovering Forbidden Topics in Language Models
by: Rager, Can, et al.
Published: (2025)
by: Rager, Can, et al.
Published: (2025)
Soft Best-of-n Sampling for Model Alignment
by: Verdun, Claudio Mayrink, et al.
Published: (2025)
by: Verdun, Claudio Mayrink, et al.
Published: (2025)
Automatically Finding Rule-Based Neurons in OthelloGPT
by: Singh, Aditya, et al.
Published: (2025)
by: Singh, Aditya, et al.
Published: (2025)
Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability
by: Bhalla, Usha, et al.
Published: (2025)
by: Bhalla, Usha, et al.
Published: (2025)
ProofCompass: Enhancing Specialized Provers with LLM Guidance
by: Wischermann, Nicolas, et al.
Published: (2025)
by: Wischermann, Nicolas, et al.
Published: (2025)
SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability
by: Karvonen, Adam, et al.
Published: (2025)
by: Karvonen, Adam, et al.
Published: (2025)
In-Context Algebra
by: Todd, Eric, et al.
Published: (2025)
by: Todd, Eric, et al.
Published: (2025)
Inference-Time Reward Hacking in Large Language Models
by: Khalaf, Hadi, et al.
Published: (2025)
by: Khalaf, Hadi, et al.
Published: (2025)
GradPCA: Leveraging NTK Alignment for Reliable Out-of-Distribution Detection
by: Seleznova, Mariia, et al.
Published: (2025)
by: Seleznova, Mariia, et al.
Published: (2025)
Polynomial Precision Dependence Solutions to Alignment Research Center Matrix Completion Problems
by: Angell, Rico
Published: (2024)
by: Angell, Rico
Published: (2024)
Multi-Group Proportional Representation for Text-to-Image Models
by: Jung, Sangwon, et al.
Published: (2025)
by: Jung, Sangwon, et al.
Published: (2025)
Emergent World Models and Latent Variable Estimation in Chess-Playing Language Models
by: Karvonen, Adam
Published: (2024)
by: Karvonen, Adam
Published: (2024)
High-Dimensional Confidence Regions in Sparse MRI
by: Hoppe, Frederik, et al.
Published: (2024)
by: Hoppe, Frederik, et al.
Published: (2024)
Non-Asymptotic Uncertainty Quantification in High-Dimensional Learning
by: Hoppe, Frederik, et al.
Published: (2024)
by: Hoppe, Frederik, et al.
Published: (2024)
Optimized Couplings for Watermarking Large Language Models
by: Tsur, Dor, et al.
Published: (2025)
by: Tsur, Dor, et al.
Published: (2025)
With or Without Replacement? Improving Confidence in Fourier Imaging
by: Hoppe, Frederik, et al.
Published: (2024)
by: Hoppe, Frederik, et al.
Published: (2024)
Erasing Conceptual Knowledge from Language Models
by: Gandikota, Rohit, et al.
Published: (2024)
by: Gandikota, Rohit, et al.
Published: (2024)
NNsight and NDIF: Democratizing Access to Open-Weight Foundation Model Internals
by: Fiotto-Kaufman, Jaden, et al.
Published: (2024)
by: Fiotto-Kaufman, Jaden, et al.
Published: (2024)
In-Context Learning Without Copying
by: Sahin, Kerem, et al.
Published: (2025)
by: Sahin, Kerem, et al.
Published: (2025)
Revisiting End-To-End Sparse Autoencoder Training: A Short Finetune Is All You Need
by: Karvonen, Adam
Published: (2025)
by: Karvonen, Adam
Published: (2025)
Get rid of your constraints and reparametrize: A study in NNLS and implicit bias
by: Chou, Hung-Hsu, et al.
Published: (2022)
by: Chou, Hung-Hsu, et al.
Published: (2022)
Fast, Scalable, Warm-Start Semidefinite Programming with Spectral Bundling and Sketching
by: Angell, Rico, et al.
Published: (2023)
by: Angell, Rico, et al.
Published: (2023)
Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning
by: Casademunt, Helena, et al.
Published: (2025)
by: Casademunt, Helena, et al.
Published: (2025)
Steering Language Models in Multi-Token Generation: A Case Study on Tense and Aspect
by: Klerings, Alina, et al.
Published: (2025)
by: Klerings, Alina, et al.
Published: (2025)
Large Language Models Share Representations of Latent Grammatical Concepts Across Typologically Diverse Languages
by: Brinkmann, Jannik, et al.
Published: (2025)
by: Brinkmann, Jannik, et al.
Published: (2025)
Mechanisms of AI Protein Folding in ESMFold
by: Lu, Kevin, et al.
Published: (2026)
by: Lu, Kevin, et al.
Published: (2026)
A Dictionary of Closed-Form Kernel Mean Embeddings
by: Briol, François-Xavier, et al.
Published: (2025)
by: Briol, François-Xavier, et al.
Published: (2025)
Is It Thinking or Cheating? Detecting Implicit Reward Hacking by Measuring Reasoning Effort
by: Wang, Xinpeng, et al.
Published: (2025)
by: Wang, Xinpeng, et al.
Published: (2025)
GOV-REK: Governed Reward Engineering Kernels for Designing Robust Multi-Agent Reinforcement Learning Systems
by: Rana, Ashish, et al.
Published: (2024)
by: Rana, Ashish, et al.
Published: (2024)
NSA: Neuro-symbolic ARC Challenge
by: Batorski, Paweł, et al.
Published: (2025)
by: Batorski, Paweł, et al.
Published: (2025)
A Tuned Value Chain Model for University Based Public Research Organisation. Case Lut Cst.
by: Vesa Karvonen
Published: (2012)
by: Vesa Karvonen
Published: (2012)
Les diverses formes de colonisation pour les chômeurs en Autriche
by: Fritz Rager
Published: (1934)
by: Fritz Rager
Published: (1934)
Apprentice training in the Austrian metal industry
by: Fritz Rager
Published: (1923)
by: Fritz Rager
Published: (1923)
The settlement of the unemployed on the land in Austria
by: Fritz Rager
Published: (1934)
by: Fritz Rager
Published: (1934)
Similar Items
-
Evaluating Sparse Autoencoders on Targeted Concept Erasure Tasks
by: Karvonen, Adam, et al.
Published: (2024) -
Mitigating Adaptive Attacks against Reasoning Models with Activation Consistency Training
by: Shah, Avidan, et al.
Published: (2026) -
Jailbreak Transferability Emerges from Shared Representations
by: Angell, Rico, et al.
Published: (2025) -
Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
by: Marks, Samuel, et al.
Published: (2024) -
Robustly Improving LLM Fairness in Realistic Settings via Interpretability
by: Karvonen, Adam, et al.
Published: (2025)