Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers
Fuente:
arXiv
Saved in:
| Main Authors: | Karvonen, Adam, Chua, James, Dumas, Clément, Fraser-Taliente, Kit, Kantamneni, Subhash, Minder, Julian, Ong, Euan, Sharma, Arnab Sen, Wen, Daniel, Evans, Owain, Marks, Samuel |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The $T^{μν}$ of the conformal scalars
by: Fraser-Taliente, Kit, et al.
Published: (2026)
by: Fraser-Taliente, Kit, et al.
Published: (2026)
Believe It or Not: How Deeply do LLMs Believe Implanted Facts?
by: Slocum, Stewart, et al.
Published: (2025)
by: Slocum, Stewart, et al.
Published: (2025)
The Consciousness Cluster: Emergent preferences of Models that Claim to be Conscious
by: Chua, James, et al.
Published: (2026)
by: Chua, James, et al.
Published: (2026)
Narrow Finetuning Leaves Clearly Readable Traces in Activation Differences
by: Minder, Julian, et al.
Published: (2025)
by: Minder, Julian, et al.
Published: (2025)
Diffusion Models for Cayley Graphs
by: Douglas, Michael R., et al.
Published: (2025)
by: Douglas, Michael R., et al.
Published: (2025)
Language Models Use Trigonometry to Do Addition
by: Kantamneni, Subhash, et al.
Published: (2025)
by: Kantamneni, Subhash, et al.
Published: (2025)
Are DeepSeek R1 And Other Reasoning Models More Faithful?
by: Chua, James, et al.
Published: (2025)
by: Chua, James, et al.
Published: (2025)
Robustly Improving LLM Fairness in Realistic Settings via Interpretability
by: Karvonen, Adam, et al.
Published: (2025)
by: Karvonen, Adam, et al.
Published: (2025)
Overcoming Sparsity Artifacts in Crosscoders to Interpret Chat-Tuning
by: Minder, Julian, et al.
Published: (2025)
by: Minder, Julian, et al.
Published: (2025)
The sphere free energy of the vector models to order $1/N$
by: Fraser-Taliente, Ludo
Published: (2025)
by: Fraser-Taliente, Ludo
Published: (2025)
Local CFTs extremise $F$
by: Fraser-Taliente, Ludo
Published: (2026)
by: Fraser-Taliente, Ludo
Published: (2026)
Quantum field theories with many fields
by: Fraser-Taliente, Ludo
Published: (2026)
by: Fraser-Taliente, Ludo
Published: (2026)
Not So Flat Metrics
by: Fraser-Taliente, Kit, et al.
Published: (2024)
by: Fraser-Taliente, Kit, et al.
Published: (2024)
Negation Neglect: When models fail to learn negations in training
by: Mayne, Harry, et al.
Published: (2026)
by: Mayne, Harry, et al.
Published: (2026)
OptPDE: Discovering Novel Integrable Systems via AI-Human Collaboration
by: Kantamneni, Subhash, et al.
Published: (2024)
by: Kantamneni, Subhash, et al.
Published: (2024)
How Do Transformers "Do" Physics? Investigating the Simple Harmonic Oscillator
by: Kantamneni, Subhash, et al.
Published: (2024)
by: Kantamneni, Subhash, et al.
Published: (2024)
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
by: Taylor, Mia, et al.
Published: (2025)
by: Taylor, Mia, et al.
Published: (2025)
Symbolic Regression with Multimodal Large Language Models and Kolmogorov Arnold Networks
by: Harvey, Thomas R., et al.
Published: (2025)
by: Harvey, Thomas R., et al.
Published: (2025)
$F$-extremization determines certain large-$N$ CFTs
by: Fraser-Taliente, Ludo, et al.
Published: (2024)
by: Fraser-Taliente, Ludo, et al.
Published: (2024)
Melonic limits of the quartic Yukawa model and general features of melonic CFTs
by: Fraser-Taliente, Ludo, et al.
Published: (2024)
by: Fraser-Taliente, Ludo, et al.
Published: (2024)
Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data
by: Treutlein, Johannes, et al.
Published: (2024)
by: Treutlein, Johannes, et al.
Published: (2024)
Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models
by: Chua, James, et al.
Published: (2025)
by: Chua, James, et al.
Published: (2025)
Subliminal Learning: Language models transmit behavioral traits via hidden signals in data
by: Cloud, Alex, et al.
Published: (2025)
by: Cloud, Alex, et al.
Published: (2025)
Scaling Laws For Scalable Oversight
by: Engels, Joshua, et al.
Published: (2025)
by: Engels, Joshua, et al.
Published: (2025)
Tell me about yourself: LLMs are aware of their learned behaviors
by: Betley, Jan, et al.
Published: (2025)
by: Betley, Jan, et al.
Published: (2025)
Enumerating Calabi-Yau Manifolds: Placing bounds on the number of diffeomorphism classes in the Kreuzer-Skarke list
by: Chandra, Aditi, et al.
Published: (2023)
by: Chandra, Aditi, et al.
Published: (2023)
Computation of Quark Masses from String Theory
by: Constantin, Andrei, et al.
Published: (2024)
by: Constantin, Andrei, et al.
Published: (2024)
Evaluating Sparse Autoencoders on Targeted Concept Erasure Tasks
by: Karvonen, Adam, et al.
Published: (2024)
by: Karvonen, Adam, et al.
Published: (2024)
Correlation Decay for Maximum Weight Matchings on Sparse Graphs
by: Lam, Wai-Kit, et al.
Published: (2025)
by: Lam, Wai-Kit, et al.
Published: (2025)
Central Limit Theorem in Disordered Monomer-Dimer Model
by: Lam, Wai-Kit, et al.
Published: (2022)
by: Lam, Wai-Kit, et al.
Published: (2022)
Separating Tongue from Thought: Activation Patching Reveals Language-Agnostic Concept Representations in Transformers
by: Dumas, Clément, et al.
Published: (2024)
by: Dumas, Clément, et al.
Published: (2024)
Fermion Masses and Mixing in String-Inspired Models
by: Constantin, Andrei, et al.
Published: (2024)
by: Constantin, Andrei, et al.
Published: (2024)
Are Sparse Autoencoders Useful? A Case Study in Sparse Probing
by: Kantamneni, Subhash, et al.
Published: (2025)
by: Kantamneni, Subhash, et al.
Published: (2025)
Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs
by: Betley, Jan, et al.
Published: (2025)
by: Betley, Jan, et al.
Published: (2025)
Efficient and Concise Explanations for Object Detection with Gaussian-Class Activation Mapping Explainer
by: Nguyen, Quoc Khanh, et al.
Published: (2024)
by: Nguyen, Quoc Khanh, et al.
Published: (2024)
nnterp: A Standardized Interface for Mechanistic Interpretability of Transformers
by: Dumas, Clément
Published: (2025)
by: Dumas, Clément
Published: (2025)
Resiliency and Reliability in the Cloud Engineering Era: Lessons from Strategy, Intelligence and AI
by: Kantamneni, Venkata Srinivas
Published: (2026)
by: Kantamneni, Venkata Srinivas
Published: (2026)
A Nonlocal Schwinger Model
by: Fraser-Taliente, Ludovic, et al.
Published: (2024)
by: Fraser-Taliente, Ludovic, et al.
Published: (2024)
Data and the Decision-Making Process.
by: Minder, Thomas
Published: (1979)
by: Minder, Thomas
Published: (1979)
Application of Systems Analysis in Designing a New System
by: Minder, Thomas
Published: (1973)
by: Minder, Thomas
Published: (1973)
Similar Items
-
The $T^{μν}$ of the conformal scalars
by: Fraser-Taliente, Kit, et al.
Published: (2026) -
Believe It or Not: How Deeply do LLMs Believe Implanted Facts?
by: Slocum, Stewart, et al.
Published: (2025) -
The Consciousness Cluster: Emergent preferences of Models that Claim to be Conscious
by: Chua, James, et al.
Published: (2026) -
Narrow Finetuning Leaves Clearly Readable Traces in Activation Differences
by: Minder, Julian, et al.
Published: (2025) -
Diffusion Models for Cayley Graphs
by: Douglas, Michael R., et al.
Published: (2025)