How does GPT-2 Predict Acronyms? Extracting and Understanding a Circuit via Mechanistic Interpretability
Fuente:
arXiv
Guardado en:
| Autores principales: | García-Carrasco, Jorge, Maté, Alejandro, Trujillo, Juan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Detecting and Understanding Vulnerabilities in Language Models via Mechanistic Interpretability
por: García-Carrasco, Jorge, et al.
Publicado: (2024)
por: García-Carrasco, Jorge, et al.
Publicado: (2024)
Extracting Interpretable Task-Specific Circuits from Large Language Models for Faster Inference
por: García-Carrasco, Jorge, et al.
Publicado: (2024)
por: García-Carrasco, Jorge, et al.
Publicado: (2024)
Refining ADHD diagnosis with EEG: The impact of preprocessing and temporal segmentation on classification accuracy
por: García-Ponsoda, Sandra, et al.
Publicado: (2024)
por: García-Ponsoda, Sandra, et al.
Publicado: (2024)
Dictionary Learning Improves Patch-Free Circuit Discovery in Mechanistic Interpretability: A Case Study on Othello-GPT
por: He, Zhengfu, et al.
Publicado: (2024)
por: He, Zhengfu, et al.
Publicado: (2024)
A Mechanistic Account of Attention Sinks in GPT-2: One Circuit, Broader Implications for Mitigation
por: Ran-Milo, Yuval, et al.
Publicado: (2026)
por: Ran-Milo, Yuval, et al.
Publicado: (2026)
Mechanistic Interpretability of GPT-like Models on Summarization Tasks
por: Mishra, Anurag
Publicado: (2025)
por: Mishra, Anurag
Publicado: (2025)
MechRL: Reinforcement Learning Agents Perform Circuit Discovery for Mechanistic Interpretability
por: Khadka, Barsat
Publicado: (2026)
por: Khadka, Barsat
Publicado: (2026)
Formal Mechanistic Interpretability: Automated Circuit Discovery with Provable Guarantees
por: Hadad, Itamar, et al.
Publicado: (2026)
por: Hadad, Itamar, et al.
Publicado: (2026)
Adaptive Circuit Behavior and Generalization in Mechanistic Interpretability
por: Nainani, Jatin, et al.
Publicado: (2024)
por: Nainani, Jatin, et al.
Publicado: (2024)
Interpretable Deep Learning for Polar Mechanistic Reaction Prediction
por: Miller, Ryan J., et al.
Publicado: (2025)
por: Miller, Ryan J., et al.
Publicado: (2025)
How do Large Language Models Understand Relevance? A Mechanistic Interpretability Perspective
por: Liu, Qi, et al.
Publicado: (2025)
por: Liu, Qi, et al.
Publicado: (2025)
Unboxing the Black Box: Mechanistic Interpretability for Algorithmic Understanding of Neural Networks
por: Kowalska, Bianka, et al.
Publicado: (2025)
por: Kowalska, Bianka, et al.
Publicado: (2025)
Open Problems in Mechanistic Interpretability
por: Sharkey, Lee, et al.
Publicado: (2025)
por: Sharkey, Lee, et al.
Publicado: (2025)
Reasoning Circuits in Language Models: A Mechanistic Interpretation of Syllogistic Inference
por: Kim, Geonhee, et al.
Publicado: (2024)
por: Kim, Geonhee, et al.
Publicado: (2024)
Exemplar Partitioning for Mechanistic Interpretability
por: Rumbelow, Jessica
Publicado: (2026)
por: Rumbelow, Jessica
Publicado: (2026)
From Mechanistic to Compositional Interpretability
por: Gauderis, Ward, et al.
Publicado: (2026)
por: Gauderis, Ward, et al.
Publicado: (2026)
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability
por: Baroni, Luca, et al.
Publicado: (2025)
por: Baroni, Luca, et al.
Publicado: (2025)
Mechanistic Analysis of Circuit Preservation in Federated Learning
por: Haseeb, Muhammad, et al.
Publicado: (2025)
por: Haseeb, Muhammad, et al.
Publicado: (2025)
Enabling the Application of Graph Neural Networks on Graphs With Unknown Connectivity
por: Jorge García‐Carrasco, et al.
Publicado: (2025)
por: Jorge García‐Carrasco, et al.
Publicado: (2025)
Mechanistic Interpretability of Reinforcement Learning Agents
por: Trim, Tristan, et al.
Publicado: (2024)
por: Trim, Tristan, et al.
Publicado: (2024)
Validating Mechanistic Interpretations: An Axiomatic Approach
por: Palumbo, Nils, et al.
Publicado: (2024)
por: Palumbo, Nils, et al.
Publicado: (2024)
Mechanistic Interpretability for Neural TSP Solvers
por: Narad, Reuben, et al.
Publicado: (2025)
por: Narad, Reuben, et al.
Publicado: (2025)
Identifying a Circuit for Verb Conjugation in GPT-2
por: Africa, David Demitri
Publicado: (2025)
por: Africa, David Demitri
Publicado: (2025)
On Mechanistic Circuits for Extractive Question-Answering
por: Basu, Samyadeep, et al.
Publicado: (2025)
por: Basu, Samyadeep, et al.
Publicado: (2025)
Evaluating Brain-Inspired Modular Training in Automated Circuit Discovery for Mechanistic Interpretability
por: Nainani, Jatin
Publicado: (2024)
por: Nainani, Jatin
Publicado: (2024)
When does Self-Prediction help? Understanding Auxiliary Tasks in Reinforcement Learning
por: Voelcker, Claas, et al.
Publicado: (2024)
por: Voelcker, Claas, et al.
Publicado: (2024)
Compact Proofs of Model Performance via Mechanistic Interpretability
por: Gross, Jason, et al.
Publicado: (2024)
por: Gross, Jason, et al.
Publicado: (2024)
Using Degeneracy in the Loss Landscape for Mechanistic Interpretability
por: Bushnaq, Lucius, et al.
Publicado: (2024)
por: Bushnaq, Lucius, et al.
Publicado: (2024)
Mechanistic Interpretability of Code Correctness in LLMs via Sparse Autoencoders
por: Tahimic, Kriz, et al.
Publicado: (2025)
por: Tahimic, Kriz, et al.
Publicado: (2025)
Towards Mechanistic Interpretability of Graph Transformers via Attention Graphs
por: El, Batu, et al.
Publicado: (2025)
por: El, Batu, et al.
Publicado: (2025)
Geospatial Mechanistic Interpretability of Large Language Models
por: De Sabbata, Stef, et al.
Publicado: (2025)
por: De Sabbata, Stef, et al.
Publicado: (2025)
Dissecting and Mitigating Diffusion Bias via Mechanistic Interpretability
por: Shi, Yingdong, et al.
Publicado: (2025)
por: Shi, Yingdong, et al.
Publicado: (2025)
Bridging Mechanistic Interpretability and Prompt Engineering with Gradient Ascent for Interpretable Persona Control
por: Saini, Harshvardhan, et al.
Publicado: (2026)
por: Saini, Harshvardhan, et al.
Publicado: (2026)
Challenges in Mechanistically Interpreting Model Representations
por: Golechha, Satvik, et al.
Publicado: (2024)
por: Golechha, Satvik, et al.
Publicado: (2024)
Mechanistic Interpretability of Binary and Ternary Transformers
por: Li, Jason
Publicado: (2024)
por: Li, Jason
Publicado: (2024)
Mechanistic Interpretability of RNNs emulating Hidden Markov Models
por: Torre, Elia, et al.
Publicado: (2025)
por: Torre, Elia, et al.
Publicado: (2025)
How Do Transformers Learn to Associate Tokens: Gradient Leading Terms Bring Mechanistic Interpretability
por: Im, Shawn, et al.
Publicado: (2026)
por: Im, Shawn, et al.
Publicado: (2026)
Distilled Circuits: A Mechanistic Study of Internal Restructuring in Knowledge Distillation
por: Haskins, Reilly, et al.
Publicado: (2025)
por: Haskins, Reilly, et al.
Publicado: (2025)
Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs
por: Yan, Lecheng, et al.
Publicado: (2026)
por: Yan, Lecheng, et al.
Publicado: (2026)
Putting a Face to Forgetting: Continual Learning meets Mechanistic Interpretability
por: Masip, Sergi, et al.
Publicado: (2026)
por: Masip, Sergi, et al.
Publicado: (2026)
Ejemplares similares
-
Detecting and Understanding Vulnerabilities in Language Models via Mechanistic Interpretability
por: García-Carrasco, Jorge, et al.
Publicado: (2024) -
Extracting Interpretable Task-Specific Circuits from Large Language Models for Faster Inference
por: García-Carrasco, Jorge, et al.
Publicado: (2024) -
Refining ADHD diagnosis with EEG: The impact of preprocessing and temporal segmentation on classification accuracy
por: García-Ponsoda, Sandra, et al.
Publicado: (2024) -
Dictionary Learning Improves Patch-Free Circuit Discovery in Mechanistic Interpretability: A Case Study on Othello-GPT
por: He, Zhengfu, et al.
Publicado: (2024) -
A Mechanistic Account of Attention Sinks in GPT-2: One Circuit, Broader Implications for Mitigation
por: Ran-Milo, Yuval, et al.
Publicado: (2026)