When the Coffee Feature Activates on Coffins: An Analysis of Feature Extraction and Steering for Mechanistic Interpretability
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ronge, Raphael, Maier, Markus, Eberhardt, Frederick |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Interpretable Steering of Large Language Models with Feature Guided Activation Additions
von: Soo, Samuel, et al.
Veröffentlicht: (2025)
von: Soo, Samuel, et al.
Veröffentlicht: (2025)
Lost in Aggregation: The Causal Interpretation of the IV Estimand
von: Tsao, Danielle, et al.
Veröffentlicht: (2026)
von: Tsao, Danielle, et al.
Veröffentlicht: (2026)
Analyze Feature Flow to Enhance Interpretation and Steering in Language Models
von: Laptev, Daniil, et al.
Veröffentlicht: (2025)
von: Laptev, Daniil, et al.
Veröffentlicht: (2025)
Feature Extraction and Steering for Enhanced Chain-of-Thought Reasoning in Language Models
von: Li, Zihao, et al.
Veröffentlicht: (2025)
von: Li, Zihao, et al.
Veröffentlicht: (2025)
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs
von: Song, Xiangchen, et al.
Veröffentlicht: (2025)
von: Song, Xiangchen, et al.
Veröffentlicht: (2025)
Interpretable Prediction and Feature Selection for Survival Analysis
von: Van Ness, Mike, et al.
Veröffentlicht: (2024)
von: Van Ness, Mike, et al.
Veröffentlicht: (2024)
When Are Two Networks the Same? Tensor Similarity for Mechanistic Interpretability
von: Gonzalez, ML Nissen, et al.
Veröffentlicht: (2026)
von: Gonzalez, ML Nissen, et al.
Veröffentlicht: (2026)
DSAI: Unbiased and Interpretable Latent Feature Extraction for Data-Centric AI
von: Cho, Hyowon, et al.
Veröffentlicht: (2024)
von: Cho, Hyowon, et al.
Veröffentlicht: (2024)
Mechanistic Permutability: Match Features Across Layers
von: Balagansky, Nikita, et al.
Veröffentlicht: (2024)
von: Balagansky, Nikita, et al.
Veröffentlicht: (2024)
Investigating Graph Neural Networks and Classical Feature-Extraction Techniques in Activity-Cliff and Molecular Property Prediction
von: Dablander, Markus
Veröffentlicht: (2024)
von: Dablander, Markus
Veröffentlicht: (2024)
Comparing Feature Importance and Rule Extraction for Interpretability on Text Data
von: Lopardo, Gianluigi, et al.
Veröffentlicht: (2022)
von: Lopardo, Gianluigi, et al.
Veröffentlicht: (2022)
Interpreting and Steering State-Space Models via Activation Subspace Bottlenecks
von: Mohan, Vamshi Sunku, et al.
Veröffentlicht: (2026)
von: Mohan, Vamshi Sunku, et al.
Veröffentlicht: (2026)
Controlling for discrete unmeasured confounding in nonlinear causal models
von: Burauel, Patrick, et al.
Veröffentlicht: (2024)
von: Burauel, Patrick, et al.
Veröffentlicht: (2024)
PHLP: Sole Persistent Homology for Link Prediction - Interpretable Feature Extraction
von: You, Junwon, et al.
Veröffentlicht: (2024)
von: You, Junwon, et al.
Veröffentlicht: (2024)
To Steer or Not to Steer? Mechanistic Error Reduction with Abstention for Language Models
von: Hedström, Anna, et al.
Veröffentlicht: (2025)
von: Hedström, Anna, et al.
Veröffentlicht: (2025)
When a Zero-Shooter Cheats: Improving Age Estimation via Activation Steering
von: Imgrund, Erik, et al.
Veröffentlicht: (2026)
von: Imgrund, Erik, et al.
Veröffentlicht: (2026)
IDP-PGFE: An Interpretable Disruption Predictor based on Physics-Guided Feature Extraction
von: Shen, Chengshuo, et al.
Veröffentlicht: (2022)
von: Shen, Chengshuo, et al.
Veröffentlicht: (2022)
Mind the Performance Gap: Capability-Behavior Trade-offs in Feature Steering
von: Sprejer, Eitan, et al.
Veröffentlicht: (2026)
von: Sprejer, Eitan, et al.
Veröffentlicht: (2026)
Steered Generation via Gradient-Based Optimization on Sparse Query Features
von: Bhattacharyya, Sumanta, et al.
Veröffentlicht: (2026)
von: Bhattacharyya, Sumanta, et al.
Veröffentlicht: (2026)
Focus On This, Not That! Steering LLMs with Adaptive Feature Specification
von: Lamb, Tom A., et al.
Veröffentlicht: (2024)
von: Lamb, Tom A., et al.
Veröffentlicht: (2024)
Lower Bounds on the Size of Markov Equivalence Classes
von: Jahn, Erik, et al.
Veröffentlicht: (2025)
von: Jahn, Erik, et al.
Veröffentlicht: (2025)
Interpretable Features for the Assessment of Neurodegenerative Diseases through Handwriting Analysis
von: Thebaud, Thomas, et al.
Veröffentlicht: (2024)
von: Thebaud, Thomas, et al.
Veröffentlicht: (2024)
Local Feature Selection without Label or Feature Leakage for Interpretable Machine Learning Predictions
von: Oosterhuis, Harrie, et al.
Veröffentlicht: (2024)
von: Oosterhuis, Harrie, et al.
Veröffentlicht: (2024)
Feature Rivalry in Sparse Autoencoder Representations: A Mechanistic Study of Uncertainty-Driven Feature Competition in LLMs
von: Harshavardhan
Veröffentlicht: (2026)
von: Harshavardhan
Veröffentlicht: (2026)
Explaining Concept Shift with Interpretable Feature Attribution
von: Lyu, Ruiqi, et al.
Veröffentlicht: (2025)
von: Lyu, Ruiqi, et al.
Veröffentlicht: (2025)
Semantic-Guided RL for Interpretable Feature Engineering
von: Bouadi, Mohamed, et al.
Veröffentlicht: (2024)
von: Bouadi, Mohamed, et al.
Veröffentlicht: (2024)
SafeSteer: Interpretable Safety Steering with Refusal-Evasion in LLMs
von: Ghosh, Shaona, et al.
Veröffentlicht: (2025)
von: Ghosh, Shaona, et al.
Veröffentlicht: (2025)
Gradient Boosting Mapping for Dimensionality Reduction and Feature Extraction
von: Patron, Anri, et al.
Veröffentlicht: (2024)
von: Patron, Anri, et al.
Veröffentlicht: (2024)
When Can You Trust Your Explanations? A Robustness Analysis on Feature Importances
von: Vascotto, Ilaria, et al.
Veröffentlicht: (2024)
von: Vascotto, Ilaria, et al.
Veröffentlicht: (2024)
Tracking the Feature Dynamics in LLM Training: A Mechanistic Study
von: Xu, Yang, et al.
Veröffentlicht: (2024)
von: Xu, Yang, et al.
Veröffentlicht: (2024)
Stable and Interpretable Jet Physics with IRC-Safe Equivariant Feature Extraction
von: Konar, Partha, et al.
Veröffentlicht: (2025)
von: Konar, Partha, et al.
Veröffentlicht: (2025)
Exemplar Partitioning for Mechanistic Interpretability
von: Rumbelow, Jessica
Veröffentlicht: (2026)
von: Rumbelow, Jessica
Veröffentlicht: (2026)
From Mechanistic to Compositional Interpretability
von: Gauderis, Ward, et al.
Veröffentlicht: (2026)
von: Gauderis, Ward, et al.
Veröffentlicht: (2026)
Open Problems in Mechanistic Interpretability
von: Sharkey, Lee, et al.
Veröffentlicht: (2025)
von: Sharkey, Lee, et al.
Veröffentlicht: (2025)
Student sentiment Analysis Using Classification With Feature Extraction Techniques
von: Tamrakar, Latika, et al.
Veröffentlicht: (2021)
von: Tamrakar, Latika, et al.
Veröffentlicht: (2021)
Improving Steering Vectors by Targeting Sparse Autoencoder Features
von: Chalnev, Sviatoslav, et al.
Veröffentlicht: (2024)
von: Chalnev, Sviatoslav, et al.
Veröffentlicht: (2024)
SAEs Are Good for Steering -- If You Select the Right Features
von: Arad, Dana, et al.
Veröffentlicht: (2025)
von: Arad, Dana, et al.
Veröffentlicht: (2025)
When Features Beat Noise: A Feature Selection Technique Through Noise-Based Hypothesis Testing
von: Sinha, Mousam, et al.
Veröffentlicht: (2025)
von: Sinha, Mousam, et al.
Veröffentlicht: (2025)
Interpreting Emergent Features in Deep Learning-based Side-channel Analysis
von: Karayalçin, Sengim, et al.
Veröffentlicht: (2025)
von: Karayalçin, Sengim, et al.
Veröffentlicht: (2025)
Feature-Based Interpretable Surrogates for Optimization
von: Goerigk, Marc, et al.
Veröffentlicht: (2024)
von: Goerigk, Marc, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Interpretable Steering of Large Language Models with Feature Guided Activation Additions
von: Soo, Samuel, et al.
Veröffentlicht: (2025) -
Lost in Aggregation: The Causal Interpretation of the IV Estimand
von: Tsao, Danielle, et al.
Veröffentlicht: (2026) -
Analyze Feature Flow to Enhance Interpretation and Steering in Language Models
von: Laptev, Daniil, et al.
Veröffentlicht: (2025) -
Feature Extraction and Steering for Enhanced Chain-of-Thought Reasoning in Language Models
von: Li, Zihao, et al.
Veröffentlicht: (2025) -
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs
von: Song, Xiangchen, et al.
Veröffentlicht: (2025)