The Dual-Stream Transformer: Channelized Architecture for Interpretable Language Modeling
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Kerce, J. Clayton, Fox, Alexis |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Interpretable-by-Design Transformers via Architectural Stream Independence
par: Kerce, Clayton, et autres
Publié: (2026)
par: Kerce, Clayton, et autres
Publié: (2026)
Engineering Verifiable Modularity in Transformers via Per-Layer Supervision
par: Kerce, J. Clayton
Publié: (2026)
par: Kerce, J. Clayton
Publié: (2026)
Residual Stream Duality in Modern Transformer Architectures
par: Zhang, Yifan
Publié: (2026)
par: Zhang, Yifan
Publié: (2026)
GLIDR: Graph-Like Inductive Logic Programming with Differentiable Reasoning
par: Johnson, Blair, et autres
Publié: (2025)
par: Johnson, Blair, et autres
Publié: (2025)
Detecting Clinical Discrepancies in Health Coaching Agents: A Dual-Stream Memory and Reconciliation Architecture
par: Pugh, Samuel L, et autres
Publié: (2026)
par: Pugh, Samuel L, et autres
Publié: (2026)
Interpreting Key Mechanisms of Factual Recall in Transformer-Based Language Models
par: Lv, Ang, et autres
Publié: (2024)
par: Lv, Ang, et autres
Publié: (2024)
Momentum Streams for Optimizer-Inspired Transformers
par: Gai, Jingchu, et autres
Publié: (2026)
par: Gai, Jingchu, et autres
Publié: (2026)
Prototype Transformer: Towards Language Model Architectures Interpretable by Design
par: Yordanov, Yordan, et autres
Publié: (2026)
par: Yordanov, Yordan, et autres
Publié: (2026)
Variational Language Concepts for Interpreting Foundation Language Models
par: Wang, Hengyi, et autres
Publié: (2024)
par: Wang, Hengyi, et autres
Publié: (2024)
Learning to Interpret Weight Differences in Language Models
par: Goel, Avichal, et autres
Publié: (2025)
par: Goel, Avichal, et autres
Publié: (2025)
Rethinking Interpretability in the Era of Large Language Models
par: Singh, Chandan, et autres
Publié: (2024)
par: Singh, Chandan, et autres
Publié: (2024)
CriticAL: Critic Automation with Language Models
par: Li, Michael Y., et autres
Publié: (2024)
par: Li, Michael Y., et autres
Publié: (2024)
Does Transformer Interpretability Transfer to RNNs?
par: Paulo, Gonçalo, et autres
Publié: (2024)
par: Paulo, Gonçalo, et autres
Publié: (2024)
Stream separation improves Bregman conditioning in transformers
par: Kerce, James Clayton
Publié: (2026)
par: Kerce, James Clayton
Publié: (2026)
The Compositional Architecture of Regret in Large Language Models
par: Cui, Xiangxiang, et autres
Publié: (2025)
par: Cui, Xiangxiang, et autres
Publié: (2025)
Binary Autoencoder for Mechanistic Interpretability of Large Language Models
par: Cho, Hakaze, et autres
Publié: (2025)
par: Cho, Hakaze, et autres
Publié: (2025)
Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
par: Marks, Samuel, et autres
Publié: (2024)
par: Marks, Samuel, et autres
Publié: (2024)
Discovering Interpretable Algorithms by Decompiling Transformers to RASP
par: Huang, Xinting, et autres
Publié: (2026)
par: Huang, Xinting, et autres
Publié: (2026)
Episodic Memories Generation and Evaluation Benchmark for Large Language Models
par: Huet, Alexis, et autres
Publié: (2025)
par: Huet, Alexis, et autres
Publié: (2025)
Cross-Layer Discrete Concept Discovery for Interpreting Language Models
par: Garg, Ankur, et autres
Publié: (2025)
par: Garg, Ankur, et autres
Publié: (2025)
Interpreting Language Models Through Concept Descriptions: A Survey
par: Feldhus, Nils, et autres
Publié: (2025)
par: Feldhus, Nils, et autres
Publié: (2025)
SelfIE: Self-Interpretation of Large Language Model Embeddings
par: Chen, Haozhe, et autres
Publié: (2024)
par: Chen, Haozhe, et autres
Publié: (2024)
TracrBench: Generating Interpretability Testbeds with Large Language Models
par: Thurnherr, Hannes, et autres
Publié: (2024)
par: Thurnherr, Hannes, et autres
Publié: (2024)
Fine-Grained Interpretation of Political Opinions in Large Language Models
par: Hu, Jingyu, et autres
Publié: (2025)
par: Hu, Jingyu, et autres
Publié: (2025)
Superscopes: Amplifying Internal Feature Representations for Language Model Interpretation
par: Jacobi, Jonathan, et autres
Publié: (2025)
par: Jacobi, Jonathan, et autres
Publié: (2025)
Atlas-Alignment: Making Interpretability Transferable Across Language Models
par: Puri, Bruno, et autres
Publié: (2025)
par: Puri, Bruno, et autres
Publié: (2025)
Stream of Search (SoS): Learning to Search in Language
par: Gandhi, Kanishk, et autres
Publié: (2024)
par: Gandhi, Kanishk, et autres
Publié: (2024)
Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models
par: Karvonen, Adam, et autres
Publié: (2024)
par: Karvonen, Adam, et autres
Publié: (2024)
Supernova: Achieving More with Less in Transformer Architectures
par: Tanase, Andrei-Valentin, et autres
Publié: (2025)
par: Tanase, Andrei-Valentin, et autres
Publié: (2025)
Depth-Width tradeoffs in Algorithmic Reasoning of Graph Tasks with Transformers
par: Yehudai, Gilad, et autres
Publié: (2025)
par: Yehudai, Gilad, et autres
Publié: (2025)
Hymba: A Hybrid-head Architecture for Small Language Models
par: Dong, Xin, et autres
Publié: (2024)
par: Dong, Xin, et autres
Publié: (2024)
Reasoning Circuits in Language Models: A Mechanistic Interpretation of Syllogistic Inference
par: Kim, Geonhee, et autres
Publié: (2024)
par: Kim, Geonhee, et autres
Publié: (2024)
Interpretable Steering of Large Language Models with Feature Guided Activation Additions
par: Soo, Samuel, et autres
Publié: (2025)
par: Soo, Samuel, et autres
Publié: (2025)
Zero-Training Temporal Drift Detection for Transformer Sentiment Models: A Comprehensive Analysis on Authentic Social Media Streams
par: Bansal, Aayam, et autres
Publié: (2025)
par: Bansal, Aayam, et autres
Publié: (2025)
Peri-LN: Revisiting Normalization Layer in the Transformer Architecture
par: Kim, Jeonghoon, et autres
Publié: (2025)
par: Kim, Jeonghoon, et autres
Publié: (2025)
Revisiting the Shape Convention of Transformer Language Models
par: Liao, Feng-Ting, et autres
Publié: (2026)
par: Liao, Feng-Ting, et autres
Publié: (2026)
Beyond Components: Singular Vector-Based Interpretability of Transformer Circuits
par: Ahmad, Areeb, et autres
Publié: (2025)
par: Ahmad, Areeb, et autres
Publié: (2025)
Jet-Nemotron: Efficient Language Model with Post Neural Architecture Search
par: Gu, Yuxian, et autres
Publié: (2025)
par: Gu, Yuxian, et autres
Publié: (2025)
DLM-Scope: Mechanistic Interpretability of Diffusion Language Models via Sparse Autoencoders
par: Wang, Xu, et autres
Publié: (2026)
par: Wang, Xu, et autres
Publié: (2026)
TeleTables: A Benchmark for Large Language Models in Telecom Table Interpretation
par: Ezzakri, Anas, et autres
Publié: (2025)
par: Ezzakri, Anas, et autres
Publié: (2025)
Documents similaires
-
Interpretable-by-Design Transformers via Architectural Stream Independence
par: Kerce, Clayton, et autres
Publié: (2026) -
Engineering Verifiable Modularity in Transformers via Per-Layer Supervision
par: Kerce, J. Clayton
Publié: (2026) -
Residual Stream Duality in Modern Transformer Architectures
par: Zhang, Yifan
Publié: (2026) -
GLIDR: Graph-Like Inductive Logic Programming with Differentiable Reasoning
par: Johnson, Blair, et autres
Publié: (2025) -
Detecting Clinical Discrepancies in Health Coaching Agents: A Dual-Stream Memory and Reconciliation Architecture
par: Pugh, Samuel L, et autres
Publié: (2026)