Finding Interpretable Prompt-Specific Circuits in Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Franco, Gabriel, Tassis, Lucas M., Rohr, Azalea, Crovella, Mark |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sparse Attention Decomposition Applied to Circuit Tracing
by: Franco, Gabriel, et al.
Published: (2024)
by: Franco, Gabriel, et al.
Published: (2024)
Singular Vectors of Attention Heads Align with Features
by: Franco, Gabriel, et al.
Published: (2026)
by: Franco, Gabriel, et al.
Published: (2026)
Reasoning Circuits in Language Models: A Mechanistic Interpretation of Syllogistic Inference
by: Kim, Geonhee, et al.
Published: (2024)
by: Kim, Geonhee, et al.
Published: (2024)
Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
by: Marks, Samuel, et al.
Published: (2024)
by: Marks, Samuel, et al.
Published: (2024)
Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
by: Lan, Michael, et al.
Published: (2023)
by: Lan, Michael, et al.
Published: (2023)
ChipExpert: The Open-Source Integrated-Circuit-Design-Specific Large Language Model
by: Xu, Ning, et al.
Published: (2024)
by: Xu, Ning, et al.
Published: (2024)
CIRCUIT: A Benchmark for Circuit Interpretation and Reasoning Capabilities of LLMs
by: Skelic, Lejla, et al.
Published: (2025)
by: Skelic, Lejla, et al.
Published: (2025)
SafeSeek: Universal Attribution of Safety Circuits in Language Models
by: Yu, Miao, et al.
Published: (2026)
by: Yu, Miao, et al.
Published: (2026)
Calibrated Language Models and How to Find Them with Label Smoothing
by: Huang, Jerry, et al.
Published: (2025)
by: Huang, Jerry, et al.
Published: (2025)
Circuits, Features, and Heuristics in Molecular Transformers
by: Varadi, Kristof, et al.
Published: (2025)
by: Varadi, Kristof, et al.
Published: (2025)
Distributed Interpretability and Control for Large Language Models
by: Desai, Dev Arpan, et al.
Published: (2026)
by: Desai, Dev Arpan, et al.
Published: (2026)
Neural Probabilistic Circuits: Enabling Compositional and Interpretable Predictions through Logical Reasoning
by: Chen, Weixin, et al.
Published: (2025)
by: Chen, Weixin, et al.
Published: (2025)
Towards Empirical Interpretation of Internal Circuits and Properties in Grokked Transformers on Modular Polynomials
by: Furuta, Hiroki, et al.
Published: (2024)
by: Furuta, Hiroki, et al.
Published: (2024)
The Interpretability of Codebooks in Model-Based Reinforcement Learning is Limited
by: Eaton, Kenneth, et al.
Published: (2024)
by: Eaton, Kenneth, et al.
Published: (2024)
Large Language Models Are Latent Variable Models: Explaining and Finding Good Demonstrations for In-Context Learning
by: Wang, Xinyi, et al.
Published: (2023)
by: Wang, Xinyi, et al.
Published: (2023)
The Expressivity Boundary of Probabilistic Circuits: A Comparison with Large Language Models
by: Zhao, Zhiyu, et al.
Published: (2026)
by: Zhao, Zhiyu, et al.
Published: (2026)
Circuit Insights: Towards Interpretability Beyond Activations
by: Golimblevskaia, Elena, et al.
Published: (2025)
by: Golimblevskaia, Elena, et al.
Published: (2025)
Latent Domain Prompt Learning for Vision-Language Models
by: Li, Zhixing, et al.
Published: (2025)
by: Li, Zhixing, et al.
Published: (2025)
An Interpretable and Scalable Framework for Evaluating Large Language Models
by: Qu, Xinhao, et al.
Published: (2026)
by: Qu, Xinhao, et al.
Published: (2026)
A Hierarchical Language Model For Interpretable Graph Reasoning
by: Khurana, Sambhav, et al.
Published: (2024)
by: Khurana, Sambhav, et al.
Published: (2024)
Medical Interpretability and Knowledge Maps of Large Language Models
by: Marinescu, Razvan, et al.
Published: (2025)
by: Marinescu, Razvan, et al.
Published: (2025)
Interpreting Language Reward Models via Contrastive Explanations
by: Jiang, Junqi, et al.
Published: (2024)
by: Jiang, Junqi, et al.
Published: (2024)
Optimizing Prompts for Large Language Models: A Causal Approach
by: Chen, Wei, et al.
Published: (2026)
by: Chen, Wei, et al.
Published: (2026)
Interpretable Physics Reasoning and Performance Taxonomy in Vision-Language Models
by: Pawar, Pranav, et al.
Published: (2025)
by: Pawar, Pranav, et al.
Published: (2025)
Towards Interpretable Foundation Models of Robot Behavior: A Task Specific Policy Generation Approach
by: Sheidlower, Isaac, et al.
Published: (2024)
by: Sheidlower, Isaac, et al.
Published: (2024)
Context-Adaptive Multi-Prompt Embedding with Large Language Models for Vision-Language Alignment
by: Kim, Dahun, et al.
Published: (2025)
by: Kim, Dahun, et al.
Published: (2025)
KnowledgeVIS: Interpreting Language Models by Comparing Fill-in-the-Blank Prompts
by: Coscia, Adam, et al.
Published: (2024)
by: Coscia, Adam, et al.
Published: (2024)
InverseScope: Scalable Activation Inversion for Interpreting Large Language Models
by: Luo, Yifan, et al.
Published: (2025)
by: Luo, Yifan, et al.
Published: (2025)
Local Entropy Search over Descent Sequences for Bayesian Optimization
by: Stenger, David, et al.
Published: (2025)
by: Stenger, David, et al.
Published: (2025)
Beyond Components: Singular Vector-Based Interpretability of Transformer Circuits
by: Ahmad, Areeb, et al.
Published: (2025)
by: Ahmad, Areeb, et al.
Published: (2025)
Task-Circuit Quantization: Leveraging Knowledge Localization and Interpretability for Compression
by: Xiao, Hanqi, et al.
Published: (2025)
by: Xiao, Hanqi, et al.
Published: (2025)
Hard Prompts Made Interpretable: Sparse Entropy Regularization for Prompt Tuning with RL
by: Choi, Yunseon, et al.
Published: (2024)
by: Choi, Yunseon, et al.
Published: (2024)
Estimating the Probabilities of Rare Outputs in Language Models
by: Wu, Gabriel, et al.
Published: (2024)
by: Wu, Gabriel, et al.
Published: (2024)
Automatically Finding Reward Model Biases
by: Wang, Atticus, et al.
Published: (2026)
by: Wang, Atticus, et al.
Published: (2026)
AlignedCoT: Prompting Large Language Models via Native-Speaking Demonstrations
by: Yang, Zhicheng, et al.
Published: (2023)
by: Yang, Zhicheng, et al.
Published: (2023)
Enhancing Reasoning Capabilities of Small Language Models with Blueprints and Prompt Template Search
by: Han, Dongge, et al.
Published: (2025)
by: Han, Dongge, et al.
Published: (2025)
GRID: Scalable Task-Agnostic Prompt-Based Continual Learning for Language Models
by: Tiwari, Anushka, et al.
Published: (2025)
by: Tiwari, Anushka, et al.
Published: (2025)
REMA: A Unified Reasoning Manifold Framework for Interpreting Large Language Model
by: Li, Bo, et al.
Published: (2025)
by: Li, Bo, et al.
Published: (2025)
Towards Automated Semantic Interpretability in Reinforcement Learning via Vision-Language Models
by: Li, Zhaoxin, et al.
Published: (2025)
by: Li, Zhaoxin, et al.
Published: (2025)
Using Mechanistic Interpretability to Craft Adversarial Attacks against Large Language Models
by: Winninger, Thomas, et al.
Published: (2025)
by: Winninger, Thomas, et al.
Published: (2025)
Similar Items
-
Sparse Attention Decomposition Applied to Circuit Tracing
by: Franco, Gabriel, et al.
Published: (2024) -
Singular Vectors of Attention Heads Align with Features
by: Franco, Gabriel, et al.
Published: (2026) -
Reasoning Circuits in Language Models: A Mechanistic Interpretation of Syllogistic Inference
by: Kim, Geonhee, et al.
Published: (2024) -
Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
by: Marks, Samuel, et al.
Published: (2024) -
Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
by: Lan, Michael, et al.
Published: (2023)