Discovering Interpretable Algorithms by Decompiling Transformers to RASP
Fuente:
arXiv
Salvato in:
| Autori principali: | Huang, Xinting, Bakalova, Aleksandra, Bhattamishra, Satwik, Merrill, William, Hahn, Michael |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Decomposing Representation Space into Interpretable Subspaces with Unsupervised Learning
di: Huang, Xinting, et al.
Pubblicazione: (2025)
di: Huang, Xinting, et al.
Pubblicazione: (2025)
Contextualize-then-Aggregate: Circuits for In-Context Learning in Gemma-2 2B
di: Bakalova, Aleksandra, et al.
Pubblicazione: (2025)
di: Bakalova, Aleksandra, et al.
Pubblicazione: (2025)
InversionView: A General-Purpose Method for Reading Information from Neural Activations
di: Huang, Xinting, et al.
Pubblicazione: (2024)
di: Huang, Xinting, et al.
Pubblicazione: (2024)
How Few-Shot Examples Add Up: A Causal Decomposition of Function Vectors in In-Context Learning
di: Wang, Entang, et al.
Pubblicazione: (2026)
di: Wang, Entang, et al.
Pubblicazione: (2026)
Separations in the Representational Capabilities of Transformers and Recurrent Architectures
di: Bhattamishra, Satwik, et al.
Pubblicazione: (2024)
di: Bhattamishra, Satwik, et al.
Pubblicazione: (2024)
Born a Transformer -- Always a Transformer? On the Effect of Pretraining on Architectural Abilities
di: Jobanputra, Mayank, et al.
Pubblicazione: (2025)
di: Jobanputra, Mayank, et al.
Pubblicazione: (2025)
Revisiting Padded Transformer Expressivity: Which Architectural Choices Matter and Which Don't
di: Svete, Anej, et al.
Pubblicazione: (2026)
di: Svete, Anej, et al.
Pubblicazione: (2026)
Benefits and Limitations of Communication in Multi-Agent Reasoning
di: Rizvi-Martel, Michael, et al.
Pubblicazione: (2025)
di: Rizvi-Martel, Michael, et al.
Pubblicazione: (2025)
Between Circuits and Chomsky: Pre-pretraining on Formal Languages Imparts Linguistic Biases
di: Hu, Michael Y., et al.
Pubblicazione: (2025)
di: Hu, Michael Y., et al.
Pubblicazione: (2025)
Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
di: Marks, Samuel, et al.
Pubblicazione: (2024)
di: Marks, Samuel, et al.
Pubblicazione: (2024)
On the Ability of Transformers to Verify Plans
di: Sarrof, Yash, et al.
Pubblicazione: (2026)
di: Sarrof, Yash, et al.
Pubblicazione: (2026)
A Formal Framework for Understanding Length Generalization in Transformers
di: Huang, Xinting, et al.
Pubblicazione: (2024)
di: Huang, Xinting, et al.
Pubblicazione: (2024)
Does Transformer Interpretability Transfer to RNNs?
di: Paulo, Gonçalo, et al.
Pubblicazione: (2024)
di: Paulo, Gonçalo, et al.
Pubblicazione: (2024)
Limits of Transformer Language Models on Learning to Compose Algorithms
di: Thomm, Jonathan, et al.
Pubblicazione: (2024)
di: Thomm, Jonathan, et al.
Pubblicazione: (2024)
Can Language Models Discover Scaling Laws?
di: Lin, Haowei, et al.
Pubblicazione: (2025)
di: Lin, Haowei, et al.
Pubblicazione: (2025)
Algorithmic Capabilities of Random Transformers
di: Zhong, Ziqian, et al.
Pubblicazione: (2024)
di: Zhong, Ziqian, et al.
Pubblicazione: (2024)
Selection-p: Self-Supervised Task-Agnostic Prompt Compression for Faithfulness and Transferability
di: Chung, Tsz Ting, et al.
Pubblicazione: (2024)
di: Chung, Tsz Ting, et al.
Pubblicazione: (2024)
Neural Decompiling of Tracr Transformers
di: Thurnherr, Hannes, et al.
Pubblicazione: (2024)
di: Thurnherr, Hannes, et al.
Pubblicazione: (2024)
Neural Algorithmic Reasoning for Hypergraphs with Looped Transformers
di: Huang, Zekai, et al.
Pubblicazione: (2025)
di: Huang, Zekai, et al.
Pubblicazione: (2025)
Discovering Forbidden Topics in Language Models
di: Rager, Can, et al.
Pubblicazione: (2025)
di: Rager, Can, et al.
Pubblicazione: (2025)
The Dual-Stream Transformer: Channelized Architecture for Interpretable Language Modeling
di: Kerce, J. Clayton, et al.
Pubblicazione: (2026)
di: Kerce, J. Clayton, et al.
Pubblicazione: (2026)
Beyond Components: Singular Vector-Based Interpretability of Transformer Circuits
di: Ahmad, Areeb, et al.
Pubblicazione: (2025)
di: Ahmad, Areeb, et al.
Pubblicazione: (2025)
Learning a Decision Tree Algorithm with Transformers
di: Zhuang, Yufan, et al.
Pubblicazione: (2024)
di: Zhuang, Yufan, et al.
Pubblicazione: (2024)
Interpreting Key Mechanisms of Factual Recall in Transformer-Based Language Models
di: Lv, Ang, et al.
Pubblicazione: (2024)
di: Lv, Ang, et al.
Pubblicazione: (2024)
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions
di: Neo, Clement, et al.
Pubblicazione: (2024)
di: Neo, Clement, et al.
Pubblicazione: (2024)
Provably Learning Attention with Queries
di: Bhattamishra, Satwik, et al.
Pubblicazione: (2026)
di: Bhattamishra, Satwik, et al.
Pubblicazione: (2026)
Contrast-CAT: Contrasting Activations for Enhanced Interpretability in Transformer-based Text Classifiers
di: Han, Sungmin, et al.
Pubblicazione: (2025)
di: Han, Sungmin, et al.
Pubblicazione: (2025)
Discovering Latent Knowledge in Language Models Without Supervision
di: Burns, Collin, et al.
Pubblicazione: (2022)
di: Burns, Collin, et al.
Pubblicazione: (2022)
Discovering Knowledge-Critical Subnetworks in Pretrained Language Models
di: Bayazit, Deniz, et al.
Pubblicazione: (2023)
di: Bayazit, Deniz, et al.
Pubblicazione: (2023)
HyperDAS: Towards Automating Mechanistic Interpretability with Hypernetworks
di: Sun, Jiuding, et al.
Pubblicazione: (2025)
di: Sun, Jiuding, et al.
Pubblicazione: (2025)
Values in the Wild: Discovering and Analyzing Values in Real-World Language Model Interactions
di: Huang, Saffron, et al.
Pubblicazione: (2025)
di: Huang, Saffron, et al.
Pubblicazione: (2025)
Depth-Width tradeoffs in Algorithmic Reasoning of Graph Tasks with Transformers
di: Yehudai, Gilad, et al.
Pubblicazione: (2025)
di: Yehudai, Gilad, et al.
Pubblicazione: (2025)
DiscoverLLM: From Executing Intents to Discovering Them
di: Kim, Tae Soo, et al.
Pubblicazione: (2026)
di: Kim, Tae Soo, et al.
Pubblicazione: (2026)
RLAD: Training LLMs to Discover Abstractions for Solving Reasoning Problems
di: Qu, Yuxiao, et al.
Pubblicazione: (2025)
di: Qu, Yuxiao, et al.
Pubblicazione: (2025)
The Depth Ceiling: On the Limits of Large Language Models in Discovering Latent Planning
di: Xu, Yi, et al.
Pubblicazione: (2026)
di: Xu, Yi, et al.
Pubblicazione: (2026)
ARM: Discovering Agentic Reasoning Modules for Generalizable Multi-Agent Systems
di: Yao, Bohan, et al.
Pubblicazione: (2025)
di: Yao, Bohan, et al.
Pubblicazione: (2025)
Transformer Circuit Faithfulness Metrics are not Robust
di: Miller, Joseph, et al.
Pubblicazione: (2024)
di: Miller, Joseph, et al.
Pubblicazione: (2024)
Discovering Hierarchical Latent Capabilities of Language Models via Causal Representation Learning
di: Jin, Jikai, et al.
Pubblicazione: (2025)
di: Jin, Jikai, et al.
Pubblicazione: (2025)
Learning Self-Interpretation from Interpretability Artifacts: Training Lightweight Adapters on Vector-Label Pairs
di: Pepper, Keenan, et al.
Pubblicazione: (2026)
di: Pepper, Keenan, et al.
Pubblicazione: (2026)
MIB: A Mechanistic Interpretability Benchmark
di: Mueller, Aaron, et al.
Pubblicazione: (2025)
di: Mueller, Aaron, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Decomposing Representation Space into Interpretable Subspaces with Unsupervised Learning
di: Huang, Xinting, et al.
Pubblicazione: (2025) -
Contextualize-then-Aggregate: Circuits for In-Context Learning in Gemma-2 2B
di: Bakalova, Aleksandra, et al.
Pubblicazione: (2025) -
InversionView: A General-Purpose Method for Reading Information from Neural Activations
di: Huang, Xinting, et al.
Pubblicazione: (2024) -
How Few-Shot Examples Add Up: A Causal Decomposition of Function Vectors in In-Context Learning
di: Wang, Entang, et al.
Pubblicazione: (2026) -
Separations in the Representational Capabilities of Transformers and Recurrent Architectures
di: Bhattamishra, Satwik, et al.
Pubblicazione: (2024)