Reading Task Failure Off the Activations: A Sparse-Feature Audit of GPT-2 Small on Indirect Object Identification
Fuente:
arXiv
Saved in:
| Main Author: | Nasermoghadasi, Mahdi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Neuroplasticity and Corruption in Model Mechanisms: A Case Study Of Indirect Object Identification
by: Chhabra, Vishnu Kabir, et al.
Published: (2025)
by: Chhabra, Vishnu Kabir, et al.
Published: (2025)
Investigating the Indirect Object Identification circuit in Mamba
by: Ensign, Danielle, et al.
Published: (2024)
by: Ensign, Danielle, et al.
Published: (2024)
From Indirect Object Identification to Syllogisms: Exploring Binary Mechanisms in Transformer Circuits
by: Saraipour, Karim, et al.
Published: (2025)
by: Saraipour, Karim, et al.
Published: (2025)
Emergence of Minimal Circuits for Indirect Object Identification in Attention-Only Transformers
by: Adhikari, Rabin
Published: (2025)
by: Adhikari, Rabin
Published: (2025)
On the Relationship Between Activation Outliers and Feature Death in Sparse Autoencoders
by: Simon, Elana, et al.
Published: (2026)
by: Simon, Elana, et al.
Published: (2026)
MoRFI: Monotonic Sparse Autoencoder Feature Identification
by: Dimakopoulos, Dimitris, et al.
Published: (2026)
by: Dimakopoulos, Dimitris, et al.
Published: (2026)
S$^2$GPT-PINNs: Sparse and Small models for PDEs
by: Ji, Yajie, et al.
Published: (2025)
by: Ji, Yajie, et al.
Published: (2025)
AbsTopK: Rethinking Sparse Autoencoders For Bidirectional Features
by: Zhu, Xudong, et al.
Published: (2025)
by: Zhu, Xudong, et al.
Published: (2025)
Beyond Activation Patterns: A Weight-Based Out-of-Context Explanation of Sparse Autoencoder Features
by: Liu, Yiting, et al.
Published: (2026)
by: Liu, Yiting, et al.
Published: (2026)
Evaluating Open-Source Sparse Autoencoders on Disentangling Factual Knowledge in GPT-2 Small
by: Chaudhary, Maheep, et al.
Published: (2024)
by: Chaudhary, Maheep, et al.
Published: (2024)
The Sample Complexity of Membership Inference and Privacy Auditing
by: Haghifam, Mahdi, et al.
Published: (2025)
by: Haghifam, Mahdi, et al.
Published: (2025)
Evaluating Synthetic Activations composed of SAE Latents in GPT-2
by: Giglemiani, Giorgi, et al.
Published: (2024)
by: Giglemiani, Giorgi, et al.
Published: (2024)
Balance-Guided Sparse Identification of Multiscale Nonlinear PDEs with Small-coefficient Terms
by: Dang, Zhenhua, et al.
Published: (2026)
by: Dang, Zhenhua, et al.
Published: (2026)
Sparse Activations as Conformal Predictors
by: Campos, Margarida M., et al.
Published: (2025)
by: Campos, Margarida M., et al.
Published: (2025)
Indirect Attention: Turning Context Misalignment into a Feature
by: Bahaduri, Bissmella, et al.
Published: (2025)
by: Bahaduri, Bissmella, et al.
Published: (2025)
Sign-Aware Gated Sparse Autoencoders: Modeling Anticorrelated Features with Bi-Jump-ReLU Activations
by: Wieciech, Bartosz, et al.
Published: (2026)
by: Wieciech, Bartosz, et al.
Published: (2026)
Winter Soldier: Backdooring Language Models at Pre-Training with Indirect Data Poisoning
by: Bouaziz, Wassim, et al.
Published: (2025)
by: Bouaziz, Wassim, et al.
Published: (2025)
Learning Neural Networks with Sparse Activations
by: Awasthi, Pranjal, et al.
Published: (2024)
by: Awasthi, Pranjal, et al.
Published: (2024)
Differentiable Sparse Identification of Lagrangian Dynamics
by: Zhang, Zitong, et al.
Published: (2025)
by: Zhang, Zitong, et al.
Published: (2025)
Neural Activation Patterns Across Language Model Architectures: A Comprehensive Analysis of Cognitive Task Performance
by: Naser-Moghadasi, Mahdi, et al.
Published: (2026)
by: Naser-Moghadasi, Mahdi, et al.
Published: (2026)
GRAFT: Auditing Graph Neural Networks via Global Feature Attribution
by: Sahoo, Rishi Raj, et al.
Published: (2026)
by: Sahoo, Rishi Raj, et al.
Published: (2026)
Learning a Sparse Neural Network using IHT
by: Damadi, Saeed, et al.
Published: (2024)
by: Damadi, Saeed, et al.
Published: (2024)
Quantile Activation: Correcting a Failure Mode of ML Models
by: Challa, Aditya, et al.
Published: (2024)
by: Challa, Aditya, et al.
Published: (2024)
2D Stability Selection: Design Jittering for Doubly Stable Feature Selection
by: Nouraie, Mahdi, et al.
Published: (2026)
by: Nouraie, Mahdi, et al.
Published: (2026)
GPT-FT: An Efficient Automated Feature Transformation Using GPT for Sequence Reconstruction and Performance Enhancement
by: Gao, Yang, et al.
Published: (2025)
by: Gao, Yang, et al.
Published: (2025)
Fill in the Blanks: Accelerating Q-Learning with a Handful of Demonstrations in Sparse Reward Settings
by: Azad, Seyed Mahdi Basiri, et al.
Published: (2025)
by: Azad, Seyed Mahdi Basiri, et al.
Published: (2025)
Sparsely Activated Networks
by: Bizopoulos, Paschalis, et al.
Published: (2019)
by: Bizopoulos, Paschalis, et al.
Published: (2019)
SPARLING: Learning Latent Representations with Extremely Sparse Activations
by: Gupta, Kavi, et al.
Published: (2023)
by: Gupta, Kavi, et al.
Published: (2023)
AC-SINDy: Compositional Sparse Identification of Nonlinear Dynamics
by: Racioppo, Peter
Published: (2026)
by: Racioppo, Peter
Published: (2026)
Can GPT Redefine Medical Understanding? Evaluating GPT on Biomedical Machine Reading Comprehension
by: Vatsal, Shubham, et al.
Published: (2024)
by: Vatsal, Shubham, et al.
Published: (2024)
Theoretical Insights into Overparameterized Models in Multi-Task and Replay-Based Continual Learning
by: Banayeeanzade, Amin, et al.
Published: (2024)
by: Banayeeanzade, Amin, et al.
Published: (2024)
Indirectly Parameterized Concrete Autoencoders
by: Nilsson, Alfred, et al.
Published: (2024)
by: Nilsson, Alfred, et al.
Published: (2024)
ForTIFAI: Fending Off Recursive Training Induced Failure for AI Model Collapse
by: Shabgahi, Soheil Zibakhsh, et al.
Published: (2025)
by: Shabgahi, Soheil Zibakhsh, et al.
Published: (2025)
On the transferability of Sparse Autoencoders for interpreting compressed models
by: Gupte, Suchit, et al.
Published: (2025)
by: Gupte, Suchit, et al.
Published: (2025)
A Family of Adaptive Activation Functions for Mitigating Failure Modes in Physics-Informed Neural Networks
by: Murari, Krishna
Published: (2026)
by: Murari, Krishna
Published: (2026)
Q-Sparse: All Large Language Models can be Fully Sparsely-Activated
by: Wang, Hongyu, et al.
Published: (2024)
by: Wang, Hongyu, et al.
Published: (2024)
Simultaneous Identification of Sparse Structures and Communities in Heterogeneous Graphical Models
by: Shi, Dapeng, et al.
Published: (2024)
by: Shi, Dapeng, et al.
Published: (2024)
Pruning is Optimal for Learning Sparse Features in High-Dimensions
by: Vural, Nuri Mert, et al.
Published: (2024)
by: Vural, Nuri Mert, et al.
Published: (2024)
A Tighter Complexity Analysis of SparseGPT
by: Li, Xiaoyu, et al.
Published: (2024)
by: Li, Xiaoyu, et al.
Published: (2024)
DuoGPT: Training-free Dual Sparsity through Activation-aware Pruning in LLMs
by: Yin, Ruokai, et al.
Published: (2025)
by: Yin, Ruokai, et al.
Published: (2025)
Similar Items
-
Neuroplasticity and Corruption in Model Mechanisms: A Case Study Of Indirect Object Identification
by: Chhabra, Vishnu Kabir, et al.
Published: (2025) -
Investigating the Indirect Object Identification circuit in Mamba
by: Ensign, Danielle, et al.
Published: (2024) -
From Indirect Object Identification to Syllogisms: Exploring Binary Mechanisms in Transformer Circuits
by: Saraipour, Karim, et al.
Published: (2025) -
Emergence of Minimal Circuits for Indirect Object Identification in Attention-Only Transformers
by: Adhikari, Rabin
Published: (2025) -
On the Relationship Between Activation Outliers and Feature Death in Sparse Autoencoders
by: Simon, Elana, et al.
Published: (2026)