Investigating the Indirect Object Identification circuit in Mamba
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ensign, Danielle, Garriga-Alonso, Adrià |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Among Us: A Sandbox for Measuring and Detecting Agentic Deception
von: Golechha, Satvik, et al.
Veröffentlicht: (2025)
von: Golechha, Satvik, et al.
Veröffentlicht: (2025)
SynthSAEBench: Evaluating Sparse Autoencoders on Scalable Realistic Synthetic Data
von: Chanin, David, et al.
Veröffentlicht: (2026)
von: Chanin, David, et al.
Veröffentlicht: (2026)
Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders
von: Chanin, David, et al.
Veröffentlicht: (2025)
von: Chanin, David, et al.
Veröffentlicht: (2025)
Catastrophic Goodhart: regularizing RLHF with KL divergence does not mitigate heavy-tailed reward misspecification
von: Kwa, Thomas, et al.
Veröffentlicht: (2024)
von: Kwa, Thomas, et al.
Veröffentlicht: (2024)
Adversarial Circuit Evaluation
von: de Bos, Niels uit, et al.
Veröffentlicht: (2024)
von: de Bos, Niels uit, et al.
Veröffentlicht: (2024)
Feature Hedging: Correlated Features Break Narrow Sparse Autoencoders
von: Chanin, David, et al.
Veröffentlicht: (2025)
von: Chanin, David, et al.
Veröffentlicht: (2025)
InterpBench: Semi-Synthetic Transformers for Evaluating Mechanistic Interpretability Techniques
von: Gupta, Rohan, et al.
Veröffentlicht: (2024)
von: Gupta, Rohan, et al.
Veröffentlicht: (2024)
The LLM Has Left The Chat: Evidence of Bail Preferences in Large Language Models
von: Ensign, Danielle, et al.
Veröffentlicht: (2025)
von: Ensign, Danielle, et al.
Veröffentlicht: (2025)
Biases in the Blind Spot: Detecting What LLMs Fail to Mention
von: Arcuschin, Iván, et al.
Veröffentlicht: (2026)
von: Arcuschin, Iván, et al.
Veröffentlicht: (2026)
Path Channels and Plan Extension Kernels: a Mechanistic Description of Planning in a Sokoban RNN
von: Taufeeque, Mohammad, et al.
Veröffentlicht: (2025)
von: Taufeeque, Mohammad, et al.
Veröffentlicht: (2025)
Interpreting Emergent Planning in Model-Free Reinforcement Learning
von: Bush, Thomas, et al.
Veröffentlicht: (2025)
von: Bush, Thomas, et al.
Veröffentlicht: (2025)
DiFR: Inference Verification Despite Nondeterminism
von: Karvonen, Adam, et al.
Veröffentlicht: (2025)
von: Karvonen, Adam, et al.
Veröffentlicht: (2025)
Analyzing the Generalization and Reliability of Steering Vectors
von: Tan, Daniel, et al.
Veröffentlicht: (2024)
von: Tan, Daniel, et al.
Veröffentlicht: (2024)
From Indirect Object Identification to Syllogisms: Exploring Binary Mechanisms in Transformer Circuits
von: Saraipour, Karim, et al.
Veröffentlicht: (2025)
von: Saraipour, Karim, et al.
Veröffentlicht: (2025)
Neuroplasticity and Corruption in Model Mechanisms: A Case Study Of Indirect Object Identification
von: Chhabra, Vishnu Kabir, et al.
Veröffentlicht: (2025)
von: Chhabra, Vishnu Kabir, et al.
Veröffentlicht: (2025)
Emergence of Minimal Circuits for Indirect Object Identification in Attention-Only Transformers
von: Adhikari, Rabin
Veröffentlicht: (2025)
von: Adhikari, Rabin
Veröffentlicht: (2025)
Planning in a recurrent neural network that plays Sokoban
von: Taufeeque, Mohammad, et al.
Veröffentlicht: (2024)
von: Taufeeque, Mohammad, et al.
Veröffentlicht: (2024)
Reading Task Failure Off the Activations: A Sparse-Feature Audit of GPT-2 Small on Indirect Object Identification
von: Nasermoghadasi, Mahdi
Veröffentlicht: (2026)
von: Nasermoghadasi, Mahdi
Veröffentlicht: (2026)
Hypothesis Testing the Circuit Hypothesis in LLMs
von: Shi, Claudia, et al.
Veröffentlicht: (2024)
von: Shi, Claudia, et al.
Veröffentlicht: (2024)
Short-circuiting Shortcuts: Mechanistic Investigation of Shortcuts in Text Classification
von: Eshuijs, Leon, et al.
Veröffentlicht: (2025)
von: Eshuijs, Leon, et al.
Veröffentlicht: (2025)
Indirectly Parameterized Concrete Autoencoders
von: Nilsson, Alfred, et al.
Veröffentlicht: (2024)
von: Nilsson, Alfred, et al.
Veröffentlicht: (2024)
Bi-Mamba+: Bidirectional Mamba for Time Series Forecasting
von: Liang, Aobo, et al.
Veröffentlicht: (2024)
von: Liang, Aobo, et al.
Veröffentlicht: (2024)
Arbitrated Indirect Treatment Comparisons
von: Fang, Yixin, et al.
Veröffentlicht: (2025)
von: Fang, Yixin, et al.
Veröffentlicht: (2025)
Kinetic-Mamba: Mamba-Assisted Predictions of Stiff Chemical Kinetics
von: Pandey, Additi, et al.
Veröffentlicht: (2025)
von: Pandey, Additi, et al.
Veröffentlicht: (2025)
MCST-Mamba: Multivariate Mamba-Based Model for Traffic Prediction
von: Hamad, Mohamed, et al.
Veröffentlicht: (2025)
von: Hamad, Mohamed, et al.
Veröffentlicht: (2025)
Indirect Query Bayesian Optimization with Integrated Feedback
von: Zhang, Mengyan, et al.
Veröffentlicht: (2024)
von: Zhang, Mengyan, et al.
Veröffentlicht: (2024)
PackMamba: Efficient Processing of Variable-Length Sequences in Mamba training
von: Xu, Haoran, et al.
Veröffentlicht: (2024)
von: Xu, Haoran, et al.
Veröffentlicht: (2024)
Mamba Modulation: On the Length Generalization of Mamba
von: Lu, Peng, et al.
Veröffentlicht: (2025)
von: Lu, Peng, et al.
Veröffentlicht: (2025)
Targeted Sequential Indirect Experiment Design
von: Ailer, Elisabeth, et al.
Veröffentlicht: (2024)
von: Ailer, Elisabeth, et al.
Veröffentlicht: (2024)
Mamba Hawkes Process
von: Gao, Anningzhe, et al.
Veröffentlicht: (2024)
von: Gao, Anningzhe, et al.
Veröffentlicht: (2024)
TransMamba: A Sequence-Level Hybrid Transformer-Mamba Language Model
von: Li, Yixing, et al.
Veröffentlicht: (2025)
von: Li, Yixing, et al.
Veröffentlicht: (2025)
Achilles' Heel of Mamba: Essential difficulties of the Mamba architecture demonstrated by synthetic data
von: Chen, Tianyi, et al.
Veröffentlicht: (2025)
von: Chen, Tianyi, et al.
Veröffentlicht: (2025)
Practical do-Shapley Explanations with Estimand-Agnostic Causal Inference
von: Parafita, Álvaro, et al.
Veröffentlicht: (2025)
von: Parafita, Álvaro, et al.
Veröffentlicht: (2025)
Efficient Prior Calibration From Indirect Data
von: Akyildiz, O. Deniz, et al.
Veröffentlicht: (2024)
von: Akyildiz, O. Deniz, et al.
Veröffentlicht: (2024)
DeciMamba: Exploring the Length Extrapolation Potential of Mamba
von: Ben-Kish, Assaf, et al.
Veröffentlicht: (2024)
von: Ben-Kish, Assaf, et al.
Veröffentlicht: (2024)
Is Mamba Capable of In-Context Learning?
von: Grazzi, Riccardo, et al.
Veröffentlicht: (2024)
von: Grazzi, Riccardo, et al.
Veröffentlicht: (2024)
CEPAE: Conditional Entropy-Penalized Autoencoders for Time Series Counterfactuals
von: Garriga, Tomàs, et al.
Veröffentlicht: (2026)
von: Garriga, Tomàs, et al.
Veröffentlicht: (2026)
DiffuMamba: High-Throughput Diffusion LMs with Mamba Backbone
von: Singh, Vaibhav, et al.
Veröffentlicht: (2025)
von: Singh, Vaibhav, et al.
Veröffentlicht: (2025)
ms-Mamba: Multi-scale Mamba for Time-Series Forecasting
von: Karadag, Yusuf Meric, et al.
Veröffentlicht: (2025)
von: Karadag, Yusuf Meric, et al.
Veröffentlicht: (2025)
InfoMamba: An Attention-Free Hybrid Mamba-Transformer Model
von: Wang, Youjin, et al.
Veröffentlicht: (2026)
von: Wang, Youjin, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Among Us: A Sandbox for Measuring and Detecting Agentic Deception
von: Golechha, Satvik, et al.
Veröffentlicht: (2025) -
SynthSAEBench: Evaluating Sparse Autoencoders on Scalable Realistic Synthetic Data
von: Chanin, David, et al.
Veröffentlicht: (2026) -
Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders
von: Chanin, David, et al.
Veröffentlicht: (2025) -
Catastrophic Goodhart: regularizing RLHF with KL divergence does not mitigate heavy-tailed reward misspecification
von: Kwa, Thomas, et al.
Veröffentlicht: (2024) -
Adversarial Circuit Evaluation
von: de Bos, Niels uit, et al.
Veröffentlicht: (2024)