Tracking Equivalent Mechanistic Interpretations Across Neural Networks
Fuente:
arXiv
Salvato in:
| Autori principali: | Sun, Alan, Toneva, Mariya |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Perturbed examples reveal invariances shared by language models
di: Rawal, Ruchit, et al.
Pubblicazione: (2023)
di: Rawal, Ruchit, et al.
Pubblicazione: (2023)
Speech language models lack important brain-relevant semantics
di: Oota, Subba Reddy, et al.
Pubblicazione: (2023)
di: Oota, Subba Reddy, et al.
Pubblicazione: (2023)
When Language Models Lose Their Mind: The Consequences of Brain Misalignment
di: Merlin, Gabriele, et al.
Pubblicazione: (2026)
di: Merlin, Gabriele, et al.
Pubblicazione: (2026)
Brain-tuning Improves Generalizability and Efficiency of Brain Alignment in Speech Models
di: Moussa, Omer, et al.
Pubblicazione: (2025)
di: Moussa, Omer, et al.
Pubblicazione: (2025)
Mechanistic Interpretability of Binary and Ternary Transformers
di: Li, Jason
Pubblicazione: (2024)
di: Li, Jason
Pubblicazione: (2024)
Triangulation as an Acceptance Rule for Multilingual Mechanistic Interpretability
di: Long, Yanan
Pubblicazione: (2025)
di: Long, Yanan
Pubblicazione: (2025)
Language models and brains align due to more than next-word prediction and word-level information
di: Merlin, Gabriele, et al.
Pubblicazione: (2022)
di: Merlin, Gabriele, et al.
Pubblicazione: (2022)
HyperDAS: Towards Automating Mechanistic Interpretability with Hypernetworks
di: Sun, Jiuding, et al.
Pubblicazione: (2025)
di: Sun, Jiuding, et al.
Pubblicazione: (2025)
Tracking the Feature Dynamics in LLM Training: A Mechanistic Study
di: Xu, Yang, et al.
Pubblicazione: (2024)
di: Xu, Yang, et al.
Pubblicazione: (2024)
MIB: A Mechanistic Interpretability Benchmark
di: Mueller, Aaron, et al.
Pubblicazione: (2025)
di: Mueller, Aaron, et al.
Pubblicazione: (2025)
The Story is Not the Science: Execution-Grounded Evaluation of Mechanistic Interpretability Research
di: Bai, Xiaoyan, et al.
Pubblicazione: (2026)
di: Bai, Xiaoyan, et al.
Pubblicazione: (2026)
Brain-tuned Speech Models Better Reflect Speech Processing Stages in the Brain
di: Moussa, Omer, et al.
Pubblicazione: (2025)
di: Moussa, Omer, et al.
Pubblicazione: (2025)
Fine-grained Analysis of Brain-LLM Alignment through Input Attribution
di: Proietti, Michela, et al.
Pubblicazione: (2025)
di: Proietti, Michela, et al.
Pubblicazione: (2025)
Mechanistic Interpretability of GPT-like Models on Summarization Tasks
di: Mishra, Anurag
Pubblicazione: (2025)
di: Mishra, Anurag
Pubblicazione: (2025)
Binary Autoencoder for Mechanistic Interpretability of Large Language Models
di: Cho, Hakaze, et al.
Pubblicazione: (2025)
di: Cho, Hakaze, et al.
Pubblicazione: (2025)
Mechanistic Interpretability as Statistical Estimation: A Variance Analysis
di: Méloux, Maxime, et al.
Pubblicazione: (2025)
di: Méloux, Maxime, et al.
Pubblicazione: (2025)
Everything, Everywhere, All at Once: Is Mechanistic Interpretability Identifiable?
di: Méloux, Maxime, et al.
Pubblicazione: (2025)
di: Méloux, Maxime, et al.
Pubblicazione: (2025)
Finite State Automata Inside Transformers with Chain-of-Thought: A Mechanistic Study on State Tracking
di: Zhang, Yifan, et al.
Pubblicazione: (2025)
di: Zhang, Yifan, et al.
Pubblicazione: (2025)
How Do Transformers Learn to Associate Tokens: Gradient Leading Terms Bring Mechanistic Interpretability
di: Im, Shawn, et al.
Pubblicazione: (2026)
di: Im, Shawn, et al.
Pubblicazione: (2026)
Assessing Episodic Memory in LLMs with Sequence Order Recall Tasks
di: Pink, Mathis, et al.
Pubblicazione: (2024)
di: Pink, Mathis, et al.
Pubblicazione: (2024)
Missed Causes and Ambiguous Effects: Counterfactuals Pose Challenges for Interpreting Neural Networks
di: Mueller, Aaron
Pubblicazione: (2024)
di: Mueller, Aaron
Pubblicazione: (2024)
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs
di: Song, Xiangchen, et al.
Pubblicazione: (2025)
di: Song, Xiangchen, et al.
Pubblicazione: (2025)
Detecting and Understanding Vulnerabilities in Language Models via Mechanistic Interpretability
di: García-Carrasco, Jorge, et al.
Pubblicazione: (2024)
di: García-Carrasco, Jorge, et al.
Pubblicazione: (2024)
Beyond Transcription: Mechanistic Interpretability in ASR
di: Glazer, Neta, et al.
Pubblicazione: (2025)
di: Glazer, Neta, et al.
Pubblicazione: (2025)
Improving Semantic Understanding in Speech Language Models via Brain-tuning
di: Moussa, Omer, et al.
Pubblicazione: (2024)
di: Moussa, Omer, et al.
Pubblicazione: (2024)
Mechanistic Data Attribution: Tracing the Training Origins of Interpretable LLM Units
di: Chen, Jianhui, et al.
Pubblicazione: (2026)
di: Chen, Jianhui, et al.
Pubblicazione: (2026)
Reasoning Circuits in Language Models: A Mechanistic Interpretation of Syllogistic Inference
di: Kim, Geonhee, et al.
Pubblicazione: (2024)
di: Kim, Geonhee, et al.
Pubblicazione: (2024)
Transferring Linear Features Across Language Models With Model Stitching
di: Chen, Alan, et al.
Pubblicazione: (2025)
di: Chen, Alan, et al.
Pubblicazione: (2025)
DLM-Scope: Mechanistic Interpretability of Diffusion Language Models via Sparse Autoencoders
di: Wang, Xu, et al.
Pubblicazione: (2026)
di: Wang, Xu, et al.
Pubblicazione: (2026)
Intrinsic Self-Correction in LLMs: Towards Explainable Prompting via Mechanistic Interpretability
di: Lee, Yu-Ting, et al.
Pubblicazione: (2025)
di: Lee, Yu-Ting, et al.
Pubblicazione: (2025)
Mechanistic Unlearning: Robust Knowledge Unlearning and Editing via Mechanistic Localization
di: Guo, Phillip, et al.
Pubblicazione: (2024)
di: Guo, Phillip, et al.
Pubblicazione: (2024)
Beyond Accuracy: Introducing a Symbolic-Mechanistic Approach to Interpretable Evaluation
di: Habibi, Reza, et al.
Pubblicazione: (2026)
di: Habibi, Reza, et al.
Pubblicazione: (2026)
How do Large Language Models Understand Relevance? A Mechanistic Interpretability Perspective
di: Liu, Qi, et al.
Pubblicazione: (2025)
di: Liu, Qi, et al.
Pubblicazione: (2025)
A Mathematical Philosophy of Explanations in Mechanistic Interpretability -- The Strange Science Part I.i
di: Ayonrinde, Kola, et al.
Pubblicazione: (2025)
di: Ayonrinde, Kola, et al.
Pubblicazione: (2025)
Adaptive Circuit Behavior and Generalization in Mechanistic Interpretability
di: Nainani, Jatin, et al.
Pubblicazione: (2024)
di: Nainani, Jatin, et al.
Pubblicazione: (2024)
Positional Biases Shift as Inputs Approach Context Window Limits
di: Veseli, Blerta, et al.
Pubblicazione: (2025)
di: Veseli, Blerta, et al.
Pubblicazione: (2025)
Atlas-Alignment: Making Interpretability Transferable Across Language Models
di: Puri, Bruno, et al.
Pubblicazione: (2025)
di: Puri, Bruno, et al.
Pubblicazione: (2025)
On Mechanistic Circuits for Extractive Question-Answering
di: Basu, Samyadeep, et al.
Pubblicazione: (2025)
di: Basu, Samyadeep, et al.
Pubblicazione: (2025)
Mechanistic?
di: Saphra, Naomi, et al.
Pubblicazione: (2024)
di: Saphra, Naomi, et al.
Pubblicazione: (2024)
Large Language Models as Model Organisms for Human Associative Learning
di: Kolling, Camila, et al.
Pubblicazione: (2025)
di: Kolling, Camila, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Perturbed examples reveal invariances shared by language models
di: Rawal, Ruchit, et al.
Pubblicazione: (2023) -
Speech language models lack important brain-relevant semantics
di: Oota, Subba Reddy, et al.
Pubblicazione: (2023) -
When Language Models Lose Their Mind: The Consequences of Brain Misalignment
di: Merlin, Gabriele, et al.
Pubblicazione: (2026) -
Brain-tuning Improves Generalizability and Efficiency of Brain Alignment in Speech Models
di: Moussa, Omer, et al.
Pubblicazione: (2025) -
Mechanistic Interpretability of Binary and Ternary Transformers
di: Li, Jason
Pubblicazione: (2024)