A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Rai, Daking, Zhou, Yilun, Feng, Shi, Saparov, Abulhair, Yao, Ziyu |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Mechanistic Understanding of Language Models in Syntactic Code Completion
par: Miller, Samuel, et autres
Publié: (2025)
par: Miller, Samuel, et autres
Publié: (2025)
Data-driven Circuit Discovery for Interpretability of Language Models
par: Rai, Daking, et autres
Publié: (2026)
par: Rai, Daking, et autres
Publié: (2026)
An Investigation of Neuron Activation as a Unified Lens to Explain Chain-of-Thought Eliciting Arithmetic Reasoning of LLMs
par: Rai, Daking, et autres
Publié: (2024)
par: Rai, Daking, et autres
Publié: (2024)
All for One: LLMs Solve Mental Math at the Last Token With Information Transferred From Other Tokens
par: Mamidanna, Siddarth, et autres
Publié: (2025)
par: Mamidanna, Siddarth, et autres
Publié: (2025)
Failure by Interference: Language Models Make Balanced Parentheses Errors When Faulty Mechanisms Overshadow Sound Ones
par: Rai, Daking, et autres
Publié: (2025)
par: Rai, Daking, et autres
Publié: (2025)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
par: Peters, Sydney, et autres
Publié: (2025)
par: Peters, Sydney, et autres
Publié: (2025)
Mechanistic evaluation of Transformers and state space models
par: Arora, Aryaman, et autres
Publié: (2025)
par: Arora, Aryaman, et autres
Publié: (2025)
Whether, Not Which: Mechanistic Interpretability Reveals Dissociable Affect Reception and Emotion Categorization in LLMs
par: Keeman, Michael
Publié: (2026)
par: Keeman, Michael
Publié: (2026)
Causally Grounded Mechanistic Interpretability for LLMs with Faithful Natural-Language Explanations
par: Mahale, Ajay Pravin
Publié: (2026)
par: Mahale, Ajay Pravin
Publié: (2026)
Predictive Simultaneous Interpretation: Harnessing Large Language Models for Democratizing Real-Time Multilingual Communication
par: Iida, Kurando, et autres
Publié: (2024)
par: Iida, Kurando, et autres
Publié: (2024)
Adaptive Circuit Behavior and Generalization in Mechanistic Interpretability
par: Nainani, Jatin, et autres
Publié: (2024)
par: Nainani, Jatin, et autres
Publié: (2024)
Prototype Transformer: Towards Language Model Architectures Interpretable by Design
par: Yordanov, Yordan, et autres
Publié: (2026)
par: Yordanov, Yordan, et autres
Publié: (2026)
Assessment of Transformer-Based Encoder-Decoder Model for Human-Like Summarization
par: Nair, Sindhu, et autres
Publié: (2024)
par: Nair, Sindhu, et autres
Publié: (2024)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
par: Saji, Alan, et autres
Publié: (2025)
par: Saji, Alan, et autres
Publié: (2025)
Weber's Law in Transformer Magnitude Representations: Efficient Coding, Representational Geometry, and Psychophysical Laws in Language Models
par: Cacioli, Jon-Paul
Publié: (2026)
par: Cacioli, Jon-Paul
Publié: (2026)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
par: Fadli, Samih
Publié: (2025)
par: Fadli, Samih
Publié: (2025)
How Instruction-Tuning Imparts Length Control: A Cross-Lingual Mechanistic Analysis
par: Rocchetti, Elisabetta, et autres
Publié: (2025)
par: Rocchetti, Elisabetta, et autres
Publié: (2025)
Robustness of Large Language Models to Perturbations in Text
par: Singh, Ayush, et autres
Publié: (2024)
par: Singh, Ayush, et autres
Publié: (2024)
UniHetero: Could Generation Enhance Understanding for Vision-Language-Model at Large Data Scale?
par: Chen, Fengjiao, et autres
Publié: (2025)
par: Chen, Fengjiao, et autres
Publié: (2025)
AIPsy-Affect: A Keyword-Free Clinical Stimulus Battery for Mechanistic Interpretability of Emotion in Language Models
par: Keeman, Michael
Publié: (2026)
par: Keeman, Michael
Publié: (2026)
Towards Intrinsic Interpretability of Large Language Models:A Survey of Design Principles and Architectures
par: Gao, Yutong, et autres
Publié: (2026)
par: Gao, Yutong, et autres
Publié: (2026)
Interactive-KBQA: Multi-Turn Interactions for Knowledge Base Question Answering with Large Language Models
par: Xiong, Guanming, et autres
Publié: (2024)
par: Xiong, Guanming, et autres
Publié: (2024)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
par: Oketunji, Abiodun Finbarrs
Publié: (2023)
par: Oketunji, Abiodun Finbarrs
Publié: (2023)
Large Language Model (LLM) Bias Index -- LLMBI
par: Oketunji, Abiodun Finbarrs, et autres
Publié: (2023)
par: Oketunji, Abiodun Finbarrs, et autres
Publié: (2023)
Language Models are Crossword Solvers
par: Saha, Soumadeep, et autres
Publié: (2024)
par: Saha, Soumadeep, et autres
Publié: (2024)
Super Tiny Language Models
par: Hillier, Dylan, et autres
Publié: (2024)
par: Hillier, Dylan, et autres
Publié: (2024)
Egalitarian Language Representation in Language Models: It All Begins with Tokenizers
par: Velayuthan, Menan, et autres
Publié: (2024)
par: Velayuthan, Menan, et autres
Publié: (2024)
Adaptive Focus Memory for Language Models
par: Cruz, Christopher
Publié: (2025)
par: Cruz, Christopher
Publié: (2025)
PatentGPT: A Large Language Model for Intellectual Property
par: Bai, Zilong, et autres
Publié: (2024)
par: Bai, Zilong, et autres
Publié: (2024)
UA-Legal-Bench: A Benchmark for Evaluating Large Language Models on Ukrainian Legal Reasoning
par: Ovcharov, Volodymyr
Publié: (2026)
par: Ovcharov, Volodymyr
Publié: (2026)
Integrating Emotional and Linguistic Models for Ethical Compliance in Large Language Models
par: Chang, Edward Y.
Publié: (2024)
par: Chang, Edward Y.
Publié: (2024)
Language Model Circuits Are Sparse in the Neuron Basis
par: Arora, Aryaman, et autres
Publié: (2026)
par: Arora, Aryaman, et autres
Publié: (2026)
Inference to the Best Explanation in Large Language Models
par: Dalal, Dhairya, et autres
Publié: (2024)
par: Dalal, Dhairya, et autres
Publié: (2024)
Bielik 11B v3: Multilingual Large Language Model for European Languages
par: Ociepa, Krzysztof, et autres
Publié: (2025)
par: Ociepa, Krzysztof, et autres
Publié: (2025)
Beyond Black-Box Labels: Interpretable Criteria for Diagnosing Subjective NLP Tasks
par: Rair, Nisrine, et autres
Publié: (2026)
par: Rair, Nisrine, et autres
Publié: (2026)
Machine Translation Hallucination Detection for Low and High Resource Languages using Large Language Models
par: Benkirane, Kenza, et autres
Publié: (2024)
par: Benkirane, Kenza, et autres
Publié: (2024)
Future Language Modeling from Temporal Document History
par: Li, Changmao, et autres
Publié: (2024)
par: Li, Changmao, et autres
Publié: (2024)
Streamlining Redundant Layers to Compress Large Language Models
par: Chen, Xiaodong, et autres
Publié: (2024)
par: Chen, Xiaodong, et autres
Publié: (2024)
Exploring Graph Representations of Logical Forms for Language Modeling
par: Sullivan, Michael
Publié: (2025)
par: Sullivan, Michael
Publié: (2025)
A comprehensive taxonomy of hallucinations in Large Language Models
par: Cossio, Manuel
Publié: (2025)
par: Cossio, Manuel
Publié: (2025)
Documents similaires
-
Mechanistic Understanding of Language Models in Syntactic Code Completion
par: Miller, Samuel, et autres
Publié: (2025) -
Data-driven Circuit Discovery for Interpretability of Language Models
par: Rai, Daking, et autres
Publié: (2026) -
An Investigation of Neuron Activation as a Unified Lens to Explain Chain-of-Thought Eliciting Arithmetic Reasoning of LLMs
par: Rai, Daking, et autres
Publié: (2024) -
All for One: LLMs Solve Mental Math at the Last Token With Information Transferred From Other Tokens
par: Mamidanna, Siddarth, et autres
Publié: (2025) -
Failure by Interference: Language Models Make Balanced Parentheses Errors When Faulty Mechanisms Overshadow Sound Ones
par: Rai, Daking, et autres
Publié: (2025)