Mechanistic Unveiling of Transformer Circuits: Self-Influence as a Key to Model Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Lin, Hu, Lijie, Wang, Di |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PAHQ: Accelerating Automated Circuit Discovery through Mixed-Precision Inference Optimization
by: Wang, Xinhai, et al.
Published: (2025)
by: Wang, Xinhai, et al.
Published: (2025)
Beyond Scalars: Evaluating and Understanding LLM Reasoning via Geometric Progress and Stability
by: Jiang, Xinyan, et al.
Published: (2026)
by: Jiang, Xinyan, et al.
Published: (2026)
Controlling Repetition in Protein Language Models
by: Zhang, Jiahao, et al.
Published: (2026)
by: Zhang, Jiahao, et al.
Published: (2026)
EAP-GP: Mitigating Saturation Effect in Gradient-based Automated Circuit Identification
by: Zhang, Lin, et al.
Published: (2025)
by: Zhang, Lin, et al.
Published: (2025)
Adaptive Multi-Subspace Representation Steering for Attribute Alignment in Large Language Models
by: Jiang, Xinyan, et al.
Published: (2025)
by: Jiang, Xinyan, et al.
Published: (2025)
When Modalities Conflict: How Unimodal Reasoning Uncertainty Governs Preference Dynamics in MLLMs
by: Zhang, Zhuoran, et al.
Published: (2025)
by: Zhang, Zhuoran, et al.
Published: (2025)
Reasoning Circuits in Language Models: A Mechanistic Interpretation of Syllogistic Inference
by: Kim, Geonhee, et al.
Published: (2024)
by: Kim, Geonhee, et al.
Published: (2024)
When Models Outthink Their Safety: Unveiling and Mitigating Self-Jailbreak in Large Reasoning Models
by: Mao, Yingzhi, et al.
Published: (2025)
by: Mao, Yingzhi, et al.
Published: (2025)
Counting Circuits: Mechanistic Interpretability of Visual Reasoning in Large Vision-Language Models
by: Che, Liwei, et al.
Published: (2026)
by: Che, Liwei, et al.
Published: (2026)
Seeing Through Circuits: Faithful Mechanistic Interpretability for Vision Transformers
by: Żukowska, Nina, et al.
Published: (2026)
by: Żukowska, Nina, et al.
Published: (2026)
Visual Self-Fulfilling Alignment: Shaping Safety-Oriented Personas via Threat-Related Images
by: Yang, Qishun, et al.
Published: (2026)
by: Yang, Qishun, et al.
Published: (2026)
Understanding the Dynamics of Demonstration Conflict in In-Context Learning
by: Jiao, Difan, et al.
Published: (2026)
by: Jiao, Difan, et al.
Published: (2026)
When Can Large Reasoning Models Save Thinking? Mechanistic Analysis of Behavioral Divergence in Reasoning
by: Zhu, Rongzhi, et al.
Published: (2025)
by: Zhu, Rongzhi, et al.
Published: (2025)
In-Run Data Shapley for Adam Optimizer
by: Ding, Meng, et al.
Published: (2026)
by: Ding, Meng, et al.
Published: (2026)
Towards Reasoning-Preserving Unlearning in Multimodal Large Language Models
by: Li, Hongji, et al.
Published: (2025)
by: Li, Hongji, et al.
Published: (2025)
Unveiling the Reasoning Process of Large Language Models
by: Zhang, Junjie, et al.
Published: (2026)
by: Zhang, Junjie, et al.
Published: (2026)
Evaluating Data Influence in Meta Learning
by: Ren, Chenyang, et al.
Published: (2025)
by: Ren, Chenyang, et al.
Published: (2025)
Predicting LLM Output Length via Entropy-Guided Representations
by: Xie, Huanyi, et al.
Published: (2026)
by: Xie, Huanyi, et al.
Published: (2026)
The Compositional Architecture of Regret in Large Language Models
by: Cui, Xiangxiang, et al.
Published: (2025)
by: Cui, Xiangxiang, et al.
Published: (2025)
From Reasoning to Answer: Empirical, Attention-Based and Mechanistic Insights into Distilled DeepSeek R1 Models
by: Zhang, Jue, et al.
Published: (2025)
by: Zhang, Jue, et al.
Published: (2025)
Global Evolutionary Steering: Refining Activation Steering Control via Cross-Layer Consistency
by: Jiang, Xinyan, et al.
Published: (2026)
by: Jiang, Xinyan, et al.
Published: (2026)
Certified Circuits: Stability Guarantees for Mechanistic Circuits
by: Anani, Alaa, et al.
Published: (2026)
by: Anani, Alaa, et al.
Published: (2026)
Are Your Reasoning Models Reasoning or Guessing? A Mechanistic Analysis of Hierarchical Reasoning Models
by: Ren, Zirui, et al.
Published: (2026)
by: Ren, Zirui, et al.
Published: (2026)
Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model
by: Li, Tianle, et al.
Published: (2025)
by: Li, Tianle, et al.
Published: (2025)
CircuitSeer: Mining High-Quality Data by Probing Mathematical Reasoning Circuits in LLMs
by: Wang, Shaobo, et al.
Published: (2025)
by: Wang, Shaobo, et al.
Published: (2025)
AnalogAgent: Self-Improving Analog Circuit Design Automation with LLM Agents
by: Bao, Zhixuan, et al.
Published: (2026)
by: Bao, Zhixuan, et al.
Published: (2026)
Detection and Mitigation of Hallucination in Large Reasoning Models: A Mechanistic Perspective
by: Sun, Zhongxiang, et al.
Published: (2025)
by: Sun, Zhongxiang, et al.
Published: (2025)
Improving Interpretation Faithfulness for Vision Transformers
by: Hu, Lijie, et al.
Published: (2023)
by: Hu, Lijie, et al.
Published: (2023)
Uncovering Graph Reasoning in Decoder-only Transformers with Circuit Tracing
by: Dai, Xinnan, et al.
Published: (2025)
by: Dai, Xinnan, et al.
Published: (2025)
Private Language Models via Truncated Laplacian Mechanism
by: Huang, Tianhao, et al.
Published: (2024)
by: Huang, Tianhao, et al.
Published: (2024)
Faithful Interpretation for Graph Neural Networks
by: Hu, Lijie, et al.
Published: (2024)
by: Hu, Lijie, et al.
Published: (2024)
Toward Mechanistic Explanation of Deductive Reasoning in Language Models
by: Maltoni, Davide, et al.
Published: (2025)
by: Maltoni, Davide, et al.
Published: (2025)
A Mechanistic Analysis of Looped Reasoning Language Models
by: Blayney, Hugh, et al.
Published: (2026)
by: Blayney, Hugh, et al.
Published: (2026)
Dissecting Representation Misalignment in Contrastive Learning via Influence Function
by: Hu, Lijie, et al.
Published: (2024)
by: Hu, Lijie, et al.
Published: (2024)
Editable Concept Bottleneck Models
by: Hu, Lijie, et al.
Published: (2024)
by: Hu, Lijie, et al.
Published: (2024)
From Self-Attention to Markov Models: Unveiling the Dynamics of Generative Transformers
by: Ildiz, M. Emrullah, et al.
Published: (2024)
by: Ildiz, M. Emrullah, et al.
Published: (2024)
MINAR: Mechanistic Interpretability for Neural Algorithmic Reasoning
by: He, Jesse, et al.
Published: (2026)
by: He, Jesse, et al.
Published: (2026)
Skill Path: Unveiling Language Skills from Circuit Graphs
by: Chen, Hang, et al.
Published: (2024)
by: Chen, Hang, et al.
Published: (2024)
PIXEL: Adaptive Steering Via Position-wise Injection with eXact Estimated Levels under Subspace Calibration
by: Yu, Manjiang, et al.
Published: (2025)
by: Yu, Manjiang, et al.
Published: (2025)
Self-supervised Hierarchical Visual Reasoning with World Model
by: Xu, Yuanfei, et al.
Published: (2026)
by: Xu, Yuanfei, et al.
Published: (2026)
Similar Items
-
PAHQ: Accelerating Automated Circuit Discovery through Mixed-Precision Inference Optimization
by: Wang, Xinhai, et al.
Published: (2025) -
Beyond Scalars: Evaluating and Understanding LLM Reasoning via Geometric Progress and Stability
by: Jiang, Xinyan, et al.
Published: (2026) -
Controlling Repetition in Protein Language Models
by: Zhang, Jiahao, et al.
Published: (2026) -
EAP-GP: Mitigating Saturation Effect in Gradient-based Automated Circuit Identification
by: Zhang, Lin, et al.
Published: (2025) -
Adaptive Multi-Subspace Representation Steering for Attribute Alignment in Large Language Models
by: Jiang, Xinyan, et al.
Published: (2025)