Finding Interpretable Prompt-Specific Circuits in Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Franco, Gabriel, Tassis, Lucas M., Rohr, Azalea, Crovella, Mark |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Sparse Attention Decomposition Applied to Circuit Tracing
di: Franco, Gabriel, et al.
Pubblicazione: (2024)
di: Franco, Gabriel, et al.
Pubblicazione: (2024)
Singular Vectors of Attention Heads Align with Features
di: Franco, Gabriel, et al.
Pubblicazione: (2026)
di: Franco, Gabriel, et al.
Pubblicazione: (2026)
Reasoning Circuits in Language Models: A Mechanistic Interpretation of Syllogistic Inference
di: Kim, Geonhee, et al.
Pubblicazione: (2024)
di: Kim, Geonhee, et al.
Pubblicazione: (2024)
Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
di: Marks, Samuel, et al.
Pubblicazione: (2024)
di: Marks, Samuel, et al.
Pubblicazione: (2024)
Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
di: Lan, Michael, et al.
Pubblicazione: (2023)
di: Lan, Michael, et al.
Pubblicazione: (2023)
ChipExpert: The Open-Source Integrated-Circuit-Design-Specific Large Language Model
di: Xu, Ning, et al.
Pubblicazione: (2024)
di: Xu, Ning, et al.
Pubblicazione: (2024)
CIRCUIT: A Benchmark for Circuit Interpretation and Reasoning Capabilities of LLMs
di: Skelic, Lejla, et al.
Pubblicazione: (2025)
di: Skelic, Lejla, et al.
Pubblicazione: (2025)
SafeSeek: Universal Attribution of Safety Circuits in Language Models
di: Yu, Miao, et al.
Pubblicazione: (2026)
di: Yu, Miao, et al.
Pubblicazione: (2026)
Calibrated Language Models and How to Find Them with Label Smoothing
di: Huang, Jerry, et al.
Pubblicazione: (2025)
di: Huang, Jerry, et al.
Pubblicazione: (2025)
Circuits, Features, and Heuristics in Molecular Transformers
di: Varadi, Kristof, et al.
Pubblicazione: (2025)
di: Varadi, Kristof, et al.
Pubblicazione: (2025)
Distributed Interpretability and Control for Large Language Models
di: Desai, Dev Arpan, et al.
Pubblicazione: (2026)
di: Desai, Dev Arpan, et al.
Pubblicazione: (2026)
Neural Probabilistic Circuits: Enabling Compositional and Interpretable Predictions through Logical Reasoning
di: Chen, Weixin, et al.
Pubblicazione: (2025)
di: Chen, Weixin, et al.
Pubblicazione: (2025)
Towards Empirical Interpretation of Internal Circuits and Properties in Grokked Transformers on Modular Polynomials
di: Furuta, Hiroki, et al.
Pubblicazione: (2024)
di: Furuta, Hiroki, et al.
Pubblicazione: (2024)
The Interpretability of Codebooks in Model-Based Reinforcement Learning is Limited
di: Eaton, Kenneth, et al.
Pubblicazione: (2024)
di: Eaton, Kenneth, et al.
Pubblicazione: (2024)
Large Language Models Are Latent Variable Models: Explaining and Finding Good Demonstrations for In-Context Learning
di: Wang, Xinyi, et al.
Pubblicazione: (2023)
di: Wang, Xinyi, et al.
Pubblicazione: (2023)
The Expressivity Boundary of Probabilistic Circuits: A Comparison with Large Language Models
di: Zhao, Zhiyu, et al.
Pubblicazione: (2026)
di: Zhao, Zhiyu, et al.
Pubblicazione: (2026)
Circuit Insights: Towards Interpretability Beyond Activations
di: Golimblevskaia, Elena, et al.
Pubblicazione: (2025)
di: Golimblevskaia, Elena, et al.
Pubblicazione: (2025)
Latent Domain Prompt Learning for Vision-Language Models
di: Li, Zhixing, et al.
Pubblicazione: (2025)
di: Li, Zhixing, et al.
Pubblicazione: (2025)
An Interpretable and Scalable Framework for Evaluating Large Language Models
di: Qu, Xinhao, et al.
Pubblicazione: (2026)
di: Qu, Xinhao, et al.
Pubblicazione: (2026)
A Hierarchical Language Model For Interpretable Graph Reasoning
di: Khurana, Sambhav, et al.
Pubblicazione: (2024)
di: Khurana, Sambhav, et al.
Pubblicazione: (2024)
Medical Interpretability and Knowledge Maps of Large Language Models
di: Marinescu, Razvan, et al.
Pubblicazione: (2025)
di: Marinescu, Razvan, et al.
Pubblicazione: (2025)
Interpreting Language Reward Models via Contrastive Explanations
di: Jiang, Junqi, et al.
Pubblicazione: (2024)
di: Jiang, Junqi, et al.
Pubblicazione: (2024)
Optimizing Prompts for Large Language Models: A Causal Approach
di: Chen, Wei, et al.
Pubblicazione: (2026)
di: Chen, Wei, et al.
Pubblicazione: (2026)
Interpretable Physics Reasoning and Performance Taxonomy in Vision-Language Models
di: Pawar, Pranav, et al.
Pubblicazione: (2025)
di: Pawar, Pranav, et al.
Pubblicazione: (2025)
Towards Interpretable Foundation Models of Robot Behavior: A Task Specific Policy Generation Approach
di: Sheidlower, Isaac, et al.
Pubblicazione: (2024)
di: Sheidlower, Isaac, et al.
Pubblicazione: (2024)
Context-Adaptive Multi-Prompt Embedding with Large Language Models for Vision-Language Alignment
di: Kim, Dahun, et al.
Pubblicazione: (2025)
di: Kim, Dahun, et al.
Pubblicazione: (2025)
KnowledgeVIS: Interpreting Language Models by Comparing Fill-in-the-Blank Prompts
di: Coscia, Adam, et al.
Pubblicazione: (2024)
di: Coscia, Adam, et al.
Pubblicazione: (2024)
InverseScope: Scalable Activation Inversion for Interpreting Large Language Models
di: Luo, Yifan, et al.
Pubblicazione: (2025)
di: Luo, Yifan, et al.
Pubblicazione: (2025)
Local Entropy Search over Descent Sequences for Bayesian Optimization
di: Stenger, David, et al.
Pubblicazione: (2025)
di: Stenger, David, et al.
Pubblicazione: (2025)
Beyond Components: Singular Vector-Based Interpretability of Transformer Circuits
di: Ahmad, Areeb, et al.
Pubblicazione: (2025)
di: Ahmad, Areeb, et al.
Pubblicazione: (2025)
Task-Circuit Quantization: Leveraging Knowledge Localization and Interpretability for Compression
di: Xiao, Hanqi, et al.
Pubblicazione: (2025)
di: Xiao, Hanqi, et al.
Pubblicazione: (2025)
Hard Prompts Made Interpretable: Sparse Entropy Regularization for Prompt Tuning with RL
di: Choi, Yunseon, et al.
Pubblicazione: (2024)
di: Choi, Yunseon, et al.
Pubblicazione: (2024)
Estimating the Probabilities of Rare Outputs in Language Models
di: Wu, Gabriel, et al.
Pubblicazione: (2024)
di: Wu, Gabriel, et al.
Pubblicazione: (2024)
Automatically Finding Reward Model Biases
di: Wang, Atticus, et al.
Pubblicazione: (2026)
di: Wang, Atticus, et al.
Pubblicazione: (2026)
AlignedCoT: Prompting Large Language Models via Native-Speaking Demonstrations
di: Yang, Zhicheng, et al.
Pubblicazione: (2023)
di: Yang, Zhicheng, et al.
Pubblicazione: (2023)
Enhancing Reasoning Capabilities of Small Language Models with Blueprints and Prompt Template Search
di: Han, Dongge, et al.
Pubblicazione: (2025)
di: Han, Dongge, et al.
Pubblicazione: (2025)
GRID: Scalable Task-Agnostic Prompt-Based Continual Learning for Language Models
di: Tiwari, Anushka, et al.
Pubblicazione: (2025)
di: Tiwari, Anushka, et al.
Pubblicazione: (2025)
REMA: A Unified Reasoning Manifold Framework for Interpreting Large Language Model
di: Li, Bo, et al.
Pubblicazione: (2025)
di: Li, Bo, et al.
Pubblicazione: (2025)
Towards Automated Semantic Interpretability in Reinforcement Learning via Vision-Language Models
di: Li, Zhaoxin, et al.
Pubblicazione: (2025)
di: Li, Zhaoxin, et al.
Pubblicazione: (2025)
Using Mechanistic Interpretability to Craft Adversarial Attacks against Large Language Models
di: Winninger, Thomas, et al.
Pubblicazione: (2025)
di: Winninger, Thomas, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Sparse Attention Decomposition Applied to Circuit Tracing
di: Franco, Gabriel, et al.
Pubblicazione: (2024) -
Singular Vectors of Attention Heads Align with Features
di: Franco, Gabriel, et al.
Pubblicazione: (2026) -
Reasoning Circuits in Language Models: A Mechanistic Interpretation of Syllogistic Inference
di: Kim, Geonhee, et al.
Pubblicazione: (2024) -
Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
di: Marks, Samuel, et al.
Pubblicazione: (2024) -
Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
di: Lan, Michael, et al.
Pubblicazione: (2023)