Simple Mechanistic Explanations for Out-Of-Context Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Atticus, Engels, Joshua, Clive-Griffin, Oliver, Rajamanoharan, Senthooran, Nanda, Neel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Thought Branches: Interpreting LLM Reasoning Requires Resampling
von: Macar, Uzay, et al.
Veröffentlicht: (2025)
von: Macar, Uzay, et al.
Veröffentlicht: (2025)
Do I Know This Entity? Knowledge Awareness and Hallucinations in Language Models
von: Ferrando, Javier, et al.
Veröffentlicht: (2024)
von: Ferrando, Javier, et al.
Veröffentlicht: (2024)
Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning
von: Casademunt, Helena, et al.
Veröffentlicht: (2025)
von: Casademunt, Helena, et al.
Veröffentlicht: (2025)
Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
von: Arcuschin, Iván, et al.
Veröffentlicht: (2025)
von: Arcuschin, Iván, et al.
Veröffentlicht: (2025)
Are Sparse Autoencoders Useful? A Case Study in Sparse Probing
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2025)
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2025)
Towards eliciting latent knowledge from LLMs with mechanistic interpretability
von: Cywiński, Bartosz, et al.
Veröffentlicht: (2025)
von: Cywiński, Bartosz, et al.
Veröffentlicht: (2025)
Emergent Misalignment is Easy, Narrow Misalignment is Hard
von: Soligo, Anna, et al.
Veröffentlicht: (2026)
von: Soligo, Anna, et al.
Veröffentlicht: (2026)
Dense SAE Latents Are Features, Not Bugs
von: Sun, Xiaoqing, et al.
Veröffentlicht: (2025)
von: Sun, Xiaoqing, et al.
Veröffentlicht: (2025)
Convergent Linear Representations of Emergent Misalignment
von: Soligo, Anna, et al.
Veröffentlicht: (2025)
von: Soligo, Anna, et al.
Veröffentlicht: (2025)
Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
von: Lieberum, Tom, et al.
Veröffentlicht: (2024)
von: Lieberum, Tom, et al.
Veröffentlicht: (2024)
Model Organisms for Emergent Misalignment
von: Turner, Edward, et al.
Veröffentlicht: (2025)
von: Turner, Edward, et al.
Veröffentlicht: (2025)
Building Production-Ready Probes For Gemini
von: Kramár, János, et al.
Veröffentlicht: (2026)
von: Kramár, János, et al.
Veröffentlicht: (2026)
Eliciting Secret Knowledge from Language Models
von: Cywiński, Bartosz, et al.
Veröffentlicht: (2025)
von: Cywiński, Bartosz, et al.
Veröffentlicht: (2025)
Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders
von: Rajamanoharan, Senthooran, et al.
Veröffentlicht: (2024)
von: Rajamanoharan, Senthooran, et al.
Veröffentlicht: (2024)
How Well Do Models Follow Their Constitutions?
von: Jakkli, Arya, et al.
Veröffentlicht: (2026)
von: Jakkli, Arya, et al.
Veröffentlicht: (2026)
How Causal Abstraction Underpins Computational Explanation
von: Geiger, Atticus, et al.
Veröffentlicht: (2025)
von: Geiger, Atticus, et al.
Veröffentlicht: (2025)
Explorations of Self-Repair in Language Models
von: Rushing, Cody, et al.
Veröffentlicht: (2024)
von: Rushing, Cody, et al.
Veröffentlicht: (2024)
Towards Best Practices of Activation Patching in Language Models: Metrics and Methods
von: Zhang, Fred, et al.
Veröffentlicht: (2023)
von: Zhang, Fred, et al.
Veröffentlicht: (2023)
ScoNe: Benchmarking Negation Reasoning in Language Models With Fine-Tuning and In-Context Learning
von: She, Jingyuan Selena, et al.
Veröffentlicht: (2023)
von: She, Jingyuan Selena, et al.
Veröffentlicht: (2023)
Transcoders Find Interpretable LLM Feature Circuits
von: Dunefsky, Jacob, et al.
Veröffentlicht: (2024)
von: Dunefsky, Jacob, et al.
Veröffentlicht: (2024)
Summing Up the Facts: Additive Mechanisms Behind Factual Recall in LLMs
von: Chughtai, Bilal, et al.
Veröffentlicht: (2024)
von: Chughtai, Bilal, et al.
Veröffentlicht: (2024)
Thought Anchors: Which LLM Reasoning Steps Matter?
von: Bogdan, Paul C., et al.
Veröffentlicht: (2025)
von: Bogdan, Paul C., et al.
Veröffentlicht: (2025)
Out-of-Context Reasoning in Large Language Models
von: Shaki, Jonathan, et al.
Veröffentlicht: (2025)
von: Shaki, Jonathan, et al.
Veröffentlicht: (2025)
Evaluating Sparse Autoencoders on Targeted Concept Erasure Tasks
von: Karvonen, Adam, et al.
Veröffentlicht: (2024)
von: Karvonen, Adam, et al.
Veröffentlicht: (2024)
AtP*: An efficient and scalable method for localizing LLM behaviour to components
von: Kramár, János, et al.
Veröffentlicht: (2024)
von: Kramár, János, et al.
Veröffentlicht: (2024)
Generalization or Hallucination? Understanding Out-of-Context Reasoning in Transformers
von: Huang, Yixiao, et al.
Veröffentlicht: (2025)
von: Huang, Yixiao, et al.
Veröffentlicht: (2025)
Improving Dictionary Learning with Gated Sparse Autoencoders
von: Rajamanoharan, Senthooran, et al.
Veröffentlicht: (2024)
von: Rajamanoharan, Senthooran, et al.
Veröffentlicht: (2024)
Real-Time Detection of Hallucinated Entities in Long-Form Generation
von: Obeso, Oscar, et al.
Veröffentlicht: (2025)
von: Obeso, Oscar, et al.
Veröffentlicht: (2025)
HyperDAS: Towards Automating Mechanistic Interpretability with Hypernetworks
von: Sun, Jiuding, et al.
Veröffentlicht: (2025)
von: Sun, Jiuding, et al.
Veröffentlicht: (2025)
Constructing Interpretable Features from Compositional Neuron Groups
von: Shafran, Or, et al.
Veröffentlicht: (2025)
von: Shafran, Or, et al.
Veröffentlicht: (2025)
Scaling sparse feature circuit finding for in-context learning
von: Kharlapenko, Dmitrii, et al.
Veröffentlicht: (2025)
von: Kharlapenko, Dmitrii, et al.
Veröffentlicht: (2025)
Overcoming Sparsity Artifacts in Crosscoders to Interpret Chat-Tuning
von: Minder, Julian, et al.
Veröffentlicht: (2025)
von: Minder, Julian, et al.
Veröffentlicht: (2025)
What's the plan? Metrics for implicit planning in LLMs and their application to rhyme generation and question answering
von: Maar, Jim, et al.
Veröffentlicht: (2026)
von: Maar, Jim, et al.
Veröffentlicht: (2026)
Reasoning-Grounded Natural Language Explanations for Language Models
von: Cahlik, Vojtech, et al.
Veröffentlicht: (2025)
von: Cahlik, Vojtech, et al.
Veröffentlicht: (2025)
Censored LLMs as a Natural Testbed for Secret Knowledge Elicitation
von: Casademunt, Helena, et al.
Veröffentlicht: (2026)
von: Casademunt, Helena, et al.
Veröffentlicht: (2026)
Refusal Tokens: A Simple Way to Calibrate Refusals in Large Language Models
von: Jain, Neel, et al.
Veröffentlicht: (2024)
von: Jain, Neel, et al.
Veröffentlicht: (2024)
How Do Transformers Learn Variable Binding in Symbolic Programs?
von: Wu, Yiwei, et al.
Veröffentlicht: (2025)
von: Wu, Yiwei, et al.
Veröffentlicht: (2025)
A Mathematical Philosophy of Explanations in Mechanistic Interpretability -- The Strange Science Part I.i
von: Ayonrinde, Kola, et al.
Veröffentlicht: (2025)
von: Ayonrinde, Kola, et al.
Veröffentlicht: (2025)
RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations
von: Huang, Jing, et al.
Veröffentlicht: (2024)
von: Huang, Jing, et al.
Veröffentlicht: (2024)
Exploring Explanations Improves the Robustness of In-Context Learning
von: Honda, Ukyo, et al.
Veröffentlicht: (2025)
von: Honda, Ukyo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Thought Branches: Interpreting LLM Reasoning Requires Resampling
von: Macar, Uzay, et al.
Veröffentlicht: (2025) -
Do I Know This Entity? Knowledge Awareness and Hallucinations in Language Models
von: Ferrando, Javier, et al.
Veröffentlicht: (2024) -
Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning
von: Casademunt, Helena, et al.
Veröffentlicht: (2025) -
Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
von: Arcuschin, Iván, et al.
Veröffentlicht: (2025) -
Are Sparse Autoencoders Useful? A Case Study in Sparse Probing
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2025)