Bridging Mechanistic Interpretability and Prompt Engineering with Gradient Ascent for Interpretable Persona Control
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Saini, Harshvardhan, Tang, Yiming, Liu, Dianbo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Unified Theory of Sparse Dictionary Learning in Mechanistic Interpretability: Piecewise Biconvexity and Spurious Minima
von: Tang, Yiming, et al.
Veröffentlicht: (2025)
von: Tang, Yiming, et al.
Veröffentlicht: (2025)
How does My Model Fail? Automatic Identification and Interpretation of Physical Plausibility Failure Modes with Matryoshka Transcoders
von: Tang, Yiming, et al.
Veröffentlicht: (2025)
von: Tang, Yiming, et al.
Veröffentlicht: (2025)
When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models
von: Saini, Harshvardhan, et al.
Veröffentlicht: (2026)
von: Saini, Harshvardhan, et al.
Veröffentlicht: (2026)
Exemplar Partitioning for Mechanistic Interpretability
von: Rumbelow, Jessica
Veröffentlicht: (2026)
von: Rumbelow, Jessica
Veröffentlicht: (2026)
From Mechanistic to Compositional Interpretability
von: Gauderis, Ward, et al.
Veröffentlicht: (2026)
von: Gauderis, Ward, et al.
Veröffentlicht: (2026)
Open Problems in Mechanistic Interpretability
von: Sharkey, Lee, et al.
Veröffentlicht: (2025)
von: Sharkey, Lee, et al.
Veröffentlicht: (2025)
Mechanistic Interpretability for Neural TSP Solvers
von: Narad, Reuben, et al.
Veröffentlicht: (2025)
von: Narad, Reuben, et al.
Veröffentlicht: (2025)
Mechanistic Interpretability of Reinforcement Learning Agents
von: Trim, Tristan, et al.
Veröffentlicht: (2024)
von: Trim, Tristan, et al.
Veröffentlicht: (2024)
Validating Mechanistic Interpretations: An Axiomatic Approach
von: Palumbo, Nils, et al.
Veröffentlicht: (2024)
von: Palumbo, Nils, et al.
Veröffentlicht: (2024)
Is Gradient Ascent Really Necessary? Memorize to Forget for Machine Unlearning
von: Huang, Zhuo, et al.
Veröffentlicht: (2026)
von: Huang, Zhuo, et al.
Veröffentlicht: (2026)
Geospatial Mechanistic Interpretability of Large Language Models
von: De Sabbata, Stef, et al.
Veröffentlicht: (2025)
von: De Sabbata, Stef, et al.
Veröffentlicht: (2025)
Using Degeneracy in the Loss Landscape for Mechanistic Interpretability
von: Bushnaq, Lucius, et al.
Veröffentlicht: (2024)
von: Bushnaq, Lucius, et al.
Veröffentlicht: (2024)
Offline RL via Feature-Occupancy Gradient Ascent
von: Neu, Gergely, et al.
Veröffentlicht: (2024)
von: Neu, Gergely, et al.
Veröffentlicht: (2024)
Reinforcement Learning in POMDP's via Direct Gradient Ascent
von: Baxter, Jonathan, et al.
Veröffentlicht: (2025)
von: Baxter, Jonathan, et al.
Veröffentlicht: (2025)
Challenges in Mechanistically Interpreting Model Representations
von: Golechha, Satvik, et al.
Veröffentlicht: (2024)
von: Golechha, Satvik, et al.
Veröffentlicht: (2024)
Mechanistic Interpretability of Binary and Ternary Transformers
von: Li, Jason
Veröffentlicht: (2024)
von: Li, Jason
Veröffentlicht: (2024)
Intrinsic Self-Correction in LLMs: Towards Explainable Prompting via Mechanistic Interpretability
von: Lee, Yu-Ting, et al.
Veröffentlicht: (2025)
von: Lee, Yu-Ting, et al.
Veröffentlicht: (2025)
How Do Transformers Learn to Associate Tokens: Gradient Leading Terms Bring Mechanistic Interpretability
von: Im, Shawn, et al.
Veröffentlicht: (2026)
von: Im, Shawn, et al.
Veröffentlicht: (2026)
Mechanistic Interpretability of RNNs emulating Hidden Markov Models
von: Torre, Elia, et al.
Veröffentlicht: (2025)
von: Torre, Elia, et al.
Veröffentlicht: (2025)
Interpretable Deep Learning for Polar Mechanistic Reaction Prediction
von: Miller, Ryan J., et al.
Veröffentlicht: (2025)
von: Miller, Ryan J., et al.
Veröffentlicht: (2025)
Posterior Approximation using Stochastic Gradient Ascent with Adaptive Stepsize
von: Lim, Kart-Leong, et al.
Veröffentlicht: (2024)
von: Lim, Kart-Leong, et al.
Veröffentlicht: (2024)
MINAR: Mechanistic Interpretability for Neural Algorithmic Reasoning
von: He, Jesse, et al.
Veröffentlicht: (2026)
von: He, Jesse, et al.
Veröffentlicht: (2026)
Triangulation as an Acceptance Rule for Multilingual Mechanistic Interpretability
von: Long, Yanan
Veröffentlicht: (2025)
von: Long, Yanan
Veröffentlicht: (2025)
MIB: A Mechanistic Interpretability Benchmark
von: Mueller, Aaron, et al.
Veröffentlicht: (2025)
von: Mueller, Aaron, et al.
Veröffentlicht: (2025)
Label Smoothing Improves Gradient Ascent in LLM Unlearning
von: Pang, Zirui, et al.
Veröffentlicht: (2025)
von: Pang, Zirui, et al.
Veröffentlicht: (2025)
On Gradient Descent Ascent for Nonconvex-Concave Minimax Problems
von: Lin, Tianyi, et al.
Veröffentlicht: (2019)
von: Lin, Tianyi, et al.
Veröffentlicht: (2019)
When Are Two Networks the Same? Tensor Similarity for Mechanistic Interpretability
von: Gonzalez, ML Nissen, et al.
Veröffentlicht: (2026)
von: Gonzalez, ML Nissen, et al.
Veröffentlicht: (2026)
Group Equivariance Meets Mechanistic Interpretability: Equivariant Sparse Autoencoders
von: Erdogan, Ege, et al.
Veröffentlicht: (2025)
von: Erdogan, Ege, et al.
Veröffentlicht: (2025)
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs
von: Song, Xiangchen, et al.
Veröffentlicht: (2025)
von: Song, Xiangchen, et al.
Veröffentlicht: (2025)
A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models
von: Lin, Zihao, et al.
Veröffentlicht: (2025)
von: Lin, Zihao, et al.
Veröffentlicht: (2025)
Exploring Embedding Priors in Prompt-Tuning for Improved Interpretability and Control
von: Sedov, Sergey, et al.
Veröffentlicht: (2024)
von: Sedov, Sergey, et al.
Veröffentlicht: (2024)
Tracking Equivalent Mechanistic Interpretations Across Neural Networks
von: Sun, Alan, et al.
Veröffentlicht: (2026)
von: Sun, Alan, et al.
Veröffentlicht: (2026)
nnterp: A Standardized Interface for Mechanistic Interpretability of Transformers
von: Dumas, Clément
Veröffentlicht: (2025)
von: Dumas, Clément
Veröffentlicht: (2025)
Faithful and Stable Neuron Explanations for Trustworthy Mechanistic Interpretability
von: Yan, Ge, et al.
Veröffentlicht: (2025)
von: Yan, Ge, et al.
Veröffentlicht: (2025)
Mechanistic Interpretability for Transformer-based Time Series Classification
von: Kalnāre, Matīss, et al.
Veröffentlicht: (2025)
von: Kalnāre, Matīss, et al.
Veröffentlicht: (2025)
Beyond Transcription: Mechanistic Interpretability in ASR
von: Glazer, Neta, et al.
Veröffentlicht: (2025)
von: Glazer, Neta, et al.
Veröffentlicht: (2025)
Mechanistic Interpretability with Sparse Autoencoder Neural Operators
von: Tolooshams, Bahareh, et al.
Veröffentlicht: (2025)
von: Tolooshams, Bahareh, et al.
Veröffentlicht: (2025)
Putting a Face to Forgetting: Continual Learning meets Mechanistic Interpretability
von: Masip, Sergi, et al.
Veröffentlicht: (2026)
von: Masip, Sergi, et al.
Veröffentlicht: (2026)
The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability?
von: Sutter, Denis, et al.
Veröffentlicht: (2025)
von: Sutter, Denis, et al.
Veröffentlicht: (2025)
InterpBench: Semi-Synthetic Transformers for Evaluating Mechanistic Interpretability Techniques
von: Gupta, Rohan, et al.
Veröffentlicht: (2024)
von: Gupta, Rohan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
A Unified Theory of Sparse Dictionary Learning in Mechanistic Interpretability: Piecewise Biconvexity and Spurious Minima
von: Tang, Yiming, et al.
Veröffentlicht: (2025) -
How does My Model Fail? Automatic Identification and Interpretation of Physical Plausibility Failure Modes with Matryoshka Transcoders
von: Tang, Yiming, et al.
Veröffentlicht: (2025) -
When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models
von: Saini, Harshvardhan, et al.
Veröffentlicht: (2026) -
Exemplar Partitioning for Mechanistic Interpretability
von: Rumbelow, Jessica
Veröffentlicht: (2026) -
From Mechanistic to Compositional Interpretability
von: Gauderis, Ward, et al.
Veröffentlicht: (2026)