From Indirect Object Identification to Syllogisms: Exploring Binary Mechanisms in Transformer Circuits
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Saraipour, Karim, Zhang, Shichang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Emergence of Minimal Circuits for Indirect Object Identification in Attention-Only Transformers
von: Adhikari, Rabin
Veröffentlicht: (2025)
von: Adhikari, Rabin
Veröffentlicht: (2025)
How Post-Training Reshapes LLMs: A Mechanistic View on Knowledge, Truthfulness, Refusal, and Confidence
von: Du, Hongzhe, et al.
Veröffentlicht: (2025)
von: Du, Hongzhe, et al.
Veröffentlicht: (2025)
Circuit Compositions: Exploring Modular Structures in Transformer-Based Language Models
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024)
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024)
Mechanistic Interpretability of Binary and Ternary Transformers
von: Li, Jason
Veröffentlicht: (2024)
von: Li, Jason
Veröffentlicht: (2024)
Sparse Autoencoders Enable Scalable and Reliable Circuit Identification in Language Models
von: O'Neill, Charles, et al.
Veröffentlicht: (2024)
von: O'Neill, Charles, et al.
Veröffentlicht: (2024)
Categorical Syllogisms Revisited: A Review of the Logical Reasoning Abilities of LLMs for Analyzing Categorical Syllogism
von: Zong, Shi, et al.
Veröffentlicht: (2024)
von: Zong, Shi, et al.
Veröffentlicht: (2024)
Circuit Component Reuse Across Tasks in Transformer Language Models
von: Merullo, Jack, et al.
Veröffentlicht: (2023)
von: Merullo, Jack, et al.
Veröffentlicht: (2023)
Parameter-Efficient Tuning Large Language Models for Graph Representation Learning
von: Zhu, Qi, et al.
Veröffentlicht: (2024)
von: Zhu, Qi, et al.
Veröffentlicht: (2024)
Transformer Circuit Faithfulness Metrics are not Robust
von: Miller, Joseph, et al.
Veröffentlicht: (2024)
von: Miller, Joseph, et al.
Veröffentlicht: (2024)
CSP-Atlas: Concept-Specific Neural Circuits in a Sparse Python Transformer
von: Wilam, Piotr
Veröffentlicht: (2026)
von: Wilam, Piotr
Veröffentlicht: (2026)
Dissecting Multimodal In-Context Learning: Modality Asymmetries and Circuit Dynamics in modern Transformers
von: Huang, Yiran, et al.
Veröffentlicht: (2026)
von: Huang, Yiran, et al.
Veröffentlicht: (2026)
BlackboxNLP-2025 MIB Shared Task: Exploring Ensemble Strategies for Circuit Localization Methods
von: Mondorf, Philipp, et al.
Veröffentlicht: (2025)
von: Mondorf, Philipp, et al.
Veröffentlicht: (2025)
Generalized Probabilistic Attention Mechanism in Transformers
von: Heo, DongNyeong, et al.
Veröffentlicht: (2024)
von: Heo, DongNyeong, et al.
Veröffentlicht: (2024)
Judge Circuits
von: Feldhus, Nils, et al.
Veröffentlicht: (2026)
von: Feldhus, Nils, et al.
Veröffentlicht: (2026)
Beyond Components: Singular Vector-Based Interpretability of Transformer Circuits
von: Ahmad, Areeb, et al.
Veröffentlicht: (2025)
von: Ahmad, Areeb, et al.
Veröffentlicht: (2025)
Efficient Automated Circuit Discovery in Transformers using Contextual Decomposition
von: Hsu, Aliyah R., et al.
Veröffentlicht: (2024)
von: Hsu, Aliyah R., et al.
Veröffentlicht: (2024)
Reveal and Release: Iterative LLM Unlearning with Self-generated Data
von: Xie, Linxi, et al.
Veröffentlicht: (2025)
von: Xie, Linxi, et al.
Veröffentlicht: (2025)
Mechanism and Emergence of Stacked Attention Heads in Multi-Layer Transformers
von: Musat, Tiberiu
Veröffentlicht: (2024)
von: Musat, Tiberiu
Veröffentlicht: (2024)
Towards Understanding Fine-Tuning Mechanisms of LLMs via Circuit Analysis
von: Wang, Xu, et al.
Veröffentlicht: (2025)
von: Wang, Xu, et al.
Veröffentlicht: (2025)
Transformer with Controlled Attention for Synchronous Motion Captioning
von: Radouane, Karim, et al.
Veröffentlicht: (2024)
von: Radouane, Karim, et al.
Veröffentlicht: (2024)
From Sparse Dependence to Sparse Attention: Unveiling How Chain-of-Thought Enhances Transformer Sample Efficiency
von: Wen, Kaiyue, et al.
Veröffentlicht: (2024)
von: Wen, Kaiyue, et al.
Veröffentlicht: (2024)
Exploring the Impact of a Transformer's Latent Space Geometry on Downstream Task Performance
von: Marbut, Anna C., et al.
Veröffentlicht: (2024)
von: Marbut, Anna C., et al.
Veröffentlicht: (2024)
Circuit Complexity Bounds for RoPE-based Transformer Architecture
von: Chen, Bo, et al.
Veröffentlicht: (2024)
von: Chen, Bo, et al.
Veröffentlicht: (2024)
Knowledge Circuits in Pretrained Transformers
von: Yao, Yunzhi, et al.
Veröffentlicht: (2024)
von: Yao, Yunzhi, et al.
Veröffentlicht: (2024)
PathCoT: Chain-of-Thought Prompting for Zero-shot Pathology Visual Reasoning
von: Zhou, Junjie, et al.
Veröffentlicht: (2025)
von: Zhou, Junjie, et al.
Veröffentlicht: (2025)
Defending Against Indirect Prompt Injection Attacks With Spotlighting
von: Hines, Keegan, et al.
Veröffentlicht: (2024)
von: Hines, Keegan, et al.
Veröffentlicht: (2024)
Aligning Brain Activity with Advanced Transformer Models: Exploring the Role of Punctuation in Semantic Processing
von: Lamprou, Zenon, et al.
Veröffentlicht: (2025)
von: Lamprou, Zenon, et al.
Veröffentlicht: (2025)
MIPIAD: Multilingual Indirect Prompt Injection Attack Defense with Qwen -- TF-IDF Hybrid and Meta-Ensemble Learning
von: Muhtadi, Al Muhit, et al.
Veröffentlicht: (2026)
von: Muhtadi, Al Muhit, et al.
Veröffentlicht: (2026)
From Directions to Cones: Exploring Multidimensional Representations of Propositional Facts in LLMs
von: Yu, Stanley, et al.
Veröffentlicht: (2025)
von: Yu, Stanley, et al.
Veröffentlicht: (2025)
Does RLHF Scale? Exploring the Impacts From Data, Model, and Method
von: Hou, Zhenyu, et al.
Veröffentlicht: (2024)
von: Hou, Zhenyu, et al.
Veröffentlicht: (2024)
On Mechanistic Circuits for Extractive Question-Answering
von: Basu, Samyadeep, et al.
Veröffentlicht: (2025)
von: Basu, Samyadeep, et al.
Veröffentlicht: (2025)
From Interpolation to Extrapolation: Complete Length Generalization for Arithmetic Transformers
von: Duan, Shaoxiong, et al.
Veröffentlicht: (2023)
von: Duan, Shaoxiong, et al.
Veröffentlicht: (2023)
Rethinking Attention: Exploring Shallow Feed-Forward Neural Networks as an Alternative to Attention Layers in Transformers
von: Bozic, Vukasin, et al.
Veröffentlicht: (2023)
von: Bozic, Vukasin, et al.
Veröffentlicht: (2023)
Beyond the Score: Uncertainty-Calibrated LLMs for Automated Essay Assessment
von: Karim, Ahmed, et al.
Veröffentlicht: (2025)
von: Karim, Ahmed, et al.
Veröffentlicht: (2025)
Transcoders Find Interpretable LLM Feature Circuits
von: Dunefsky, Jacob, et al.
Veröffentlicht: (2024)
von: Dunefsky, Jacob, et al.
Veröffentlicht: (2024)
From Unstructured Data to In-Context Learning: Exploring What Tasks Can Be Learned and When
von: Wibisono, Kevin Christian, et al.
Veröffentlicht: (2024)
von: Wibisono, Kevin Christian, et al.
Veröffentlicht: (2024)
Measuring Uncertainty in Transformer Circuits with Effective Information Consistency
von: Krasnovsky, Anatoly A.
Veröffentlicht: (2025)
von: Krasnovsky, Anatoly A.
Veröffentlicht: (2025)
Identifying a Circuit for Verb Conjugation in GPT-2
von: Africa, David Demitri
Veröffentlicht: (2025)
von: Africa, David Demitri
Veröffentlicht: (2025)
LLM Circuit Analyses Are Consistent Across Training and Scale
von: Tigges, Curt, et al.
Veröffentlicht: (2024)
von: Tigges, Curt, et al.
Veröffentlicht: (2024)
Circuit Breaking: Removing Model Behaviors with Targeted Ablation
von: Li, Maximilian, et al.
Veröffentlicht: (2023)
von: Li, Maximilian, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Emergence of Minimal Circuits for Indirect Object Identification in Attention-Only Transformers
von: Adhikari, Rabin
Veröffentlicht: (2025) -
How Post-Training Reshapes LLMs: A Mechanistic View on Knowledge, Truthfulness, Refusal, and Confidence
von: Du, Hongzhe, et al.
Veröffentlicht: (2025) -
Circuit Compositions: Exploring Modular Structures in Transformer-Based Language Models
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024) -
Mechanistic Interpretability of Binary and Ternary Transformers
von: Li, Jason
Veröffentlicht: (2024) -
Sparse Autoencoders Enable Scalable and Reliable Circuit Identification in Language Models
von: O'Neill, Charles, et al.
Veröffentlicht: (2024)