Reasoning Circuits in Language Models: A Mechanistic Interpretation of Syllogistic Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Geonhee, Valentino, Marco, Freitas, André |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Differentiable Integer Linear Programming Solver for Explanation-Based Natural Language Inference
by: Thayaparan, Mokanarangan, et al.
Published: (2024)
by: Thayaparan, Mokanarangan, et al.
Published: (2024)
Mitigating Content Effects on Reasoning in Language Models through Fine-Grained Activation Steering
by: Valentino, Marco, et al.
Published: (2025)
by: Valentino, Marco, et al.
Published: (2025)
A Systematic Comparison of Syllogistic Reasoning in Humans and Language Models
by: Eisape, Tiwalayo, et al.
Published: (2023)
by: Eisape, Tiwalayo, et al.
Published: (2023)
Reasoning with Natural Language Explanations
by: Valentino, Marco, et al.
Published: (2024)
by: Valentino, Marco, et al.
Published: (2024)
Binary Autoencoder for Mechanistic Interpretability of Large Language Models
by: Cho, Hakaze, et al.
Published: (2025)
by: Cho, Hakaze, et al.
Published: (2025)
Eliciting Critical Reasoning in Retrieval-Augmented Language Models via Contrastive Explanations
by: Ranaldi, Leonardo, et al.
Published: (2024)
by: Ranaldi, Leonardo, et al.
Published: (2024)
Adaptive Circuit Behavior and Generalization in Mechanistic Interpretability
by: Nainani, Jatin, et al.
Published: (2024)
by: Nainani, Jatin, et al.
Published: (2024)
DLM-Scope: Mechanistic Interpretability of Diffusion Language Models via Sparse Autoencoders
by: Wang, Xu, et al.
Published: (2026)
by: Wang, Xu, et al.
Published: (2026)
Dissecting Clinical Reasoning in Language Models: A Comparative Study of Prompts and Model Adaptation Strategies
by: Jullien, Mael, et al.
Published: (2025)
by: Jullien, Mael, et al.
Published: (2025)
SemEval-2024 Task 2: Safe Biomedical Natural Language Inference for Clinical Trials
by: Jullien, Mael, et al.
Published: (2024)
by: Jullien, Mael, et al.
Published: (2024)
Revisiting In-context Learning Inference Circuit in Large Language Models
by: Cho, Hakaze, et al.
Published: (2024)
by: Cho, Hakaze, et al.
Published: (2024)
MIB: A Mechanistic Interpretability Benchmark
by: Mueller, Aaron, et al.
Published: (2025)
by: Mueller, Aaron, et al.
Published: (2025)
Mechanistic Interpretability of GPT-like Models on Summarization Tasks
by: Mishra, Anurag
Published: (2025)
by: Mishra, Anurag
Published: (2025)
Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
by: Marks, Samuel, et al.
Published: (2024)
by: Marks, Samuel, et al.
Published: (2024)
Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
by: Lan, Michael, et al.
Published: (2023)
by: Lan, Michael, et al.
Published: (2023)
Abductive Reasoning with Syllogistic Forms in Large Language Models
by: Abe, Hirohiko, et al.
Published: (2026)
by: Abe, Hirohiko, et al.
Published: (2026)
Mechanistic Interpretability as Statistical Estimation: A Variance Analysis
by: Méloux, Maxime, et al.
Published: (2025)
by: Méloux, Maxime, et al.
Published: (2025)
Inference to the Best Explanation in Large Language Models
by: Dalal, Dhairya, et al.
Published: (2024)
by: Dalal, Dhairya, et al.
Published: (2024)
SylloBio-NLI: Evaluating Large Language Models on Biomedical Syllogistic Reasoning
by: Wysocka, Magdalena, et al.
Published: (2024)
by: Wysocka, Magdalena, et al.
Published: (2024)
Towards Interpreting Language Models: A Case Study in Multi-Hop Reasoning
by: Sakarvadia, Mansi
Published: (2024)
by: Sakarvadia, Mansi
Published: (2024)
Memory Injections: Correcting Multi-Hop Reasoning Failures during Inference in Transformer-Based Language Models
by: Sakarvadia, Mansi, et al.
Published: (2023)
by: Sakarvadia, Mansi, et al.
Published: (2023)
Active Preference Inference using Language Models and Probabilistic Reasoning
by: Piriyakulkij, Wasu Top, et al.
Published: (2023)
by: Piriyakulkij, Wasu Top, et al.
Published: (2023)
HyperDAS: Towards Automating Mechanistic Interpretability with Hypernetworks
by: Sun, Jiuding, et al.
Published: (2025)
by: Sun, Jiuding, et al.
Published: (2025)
Everything, Everywhere, All at Once: Is Mechanistic Interpretability Identifiable?
by: Méloux, Maxime, et al.
Published: (2025)
by: Méloux, Maxime, et al.
Published: (2025)
Learning to Interpret Weight Differences in Language Models
by: Goel, Avichal, et al.
Published: (2025)
by: Goel, Avichal, et al.
Published: (2025)
Why Do Multilingual Reasoning Gaps Emerge in Reasoning Language Models?
by: Kang, Deokhyung, et al.
Published: (2025)
by: Kang, Deokhyung, et al.
Published: (2025)
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs
by: Song, Xiangchen, et al.
Published: (2025)
by: Song, Xiangchen, et al.
Published: (2025)
Towards Quantifying Commonsense Reasoning with Mechanistic Insights
by: Joshi, Abhinav, et al.
Published: (2025)
by: Joshi, Abhinav, et al.
Published: (2025)
Circuit Insights: Towards Interpretability Beyond Activations
by: Golimblevskaia, Elena, et al.
Published: (2025)
by: Golimblevskaia, Elena, et al.
Published: (2025)
Mechanistic Data Attribution: Tracing the Training Origins of Interpretable LLM Units
by: Chen, Jianhui, et al.
Published: (2026)
by: Chen, Jianhui, et al.
Published: (2026)
A Mathematical Philosophy of Explanations in Mechanistic Interpretability -- The Strange Science Part I.i
by: Ayonrinde, Kola, et al.
Published: (2025)
by: Ayonrinde, Kola, et al.
Published: (2025)
ECG-Reasoning-Benchmark: A Benchmark for Evaluating Clinical Reasoning Capabilities in ECG Interpretation
by: Oh, Jungwoo, et al.
Published: (2026)
by: Oh, Jungwoo, et al.
Published: (2026)
From Circuit Evidence to Mechanistic Theory: An Inductive Logic Approach
by: Aljaafari, Nura, et al.
Published: (2026)
by: Aljaafari, Nura, et al.
Published: (2026)
Block Transformer: Global-to-Local Language Modeling for Fast Inference
by: Ho, Namgyu, et al.
Published: (2024)
by: Ho, Namgyu, et al.
Published: (2024)
Intrinsic Self-Correction in LLMs: Towards Explainable Prompting via Mechanistic Interpretability
by: Lee, Yu-Ting, et al.
Published: (2025)
by: Lee, Yu-Ting, et al.
Published: (2025)
Is Inference Mediated by Distinct Semantic Structures in LLMs? A Mechanistic Interpretation
by: Aljaafari, Nura, et al.
Published: (2026)
by: Aljaafari, Nura, et al.
Published: (2026)
Self-Training Elicits Concise Reasoning in Large Language Models
by: Munkhbat, Tergel, et al.
Published: (2025)
by: Munkhbat, Tergel, et al.
Published: (2025)
Beyond Components: Singular Vector-Based Interpretability of Transformer Circuits
by: Ahmad, Areeb, et al.
Published: (2025)
by: Ahmad, Areeb, et al.
Published: (2025)
Task-Circuit Quantization: Leveraging Knowledge Localization and Interpretability for Compression
by: Xiao, Hanqi, et al.
Published: (2025)
by: Xiao, Hanqi, et al.
Published: (2025)
Graph Elicitation for Guiding Multi-Step Reasoning in Large Language Models
by: Park, Jinyoung, et al.
Published: (2023)
by: Park, Jinyoung, et al.
Published: (2023)
Similar Items
-
A Differentiable Integer Linear Programming Solver for Explanation-Based Natural Language Inference
by: Thayaparan, Mokanarangan, et al.
Published: (2024) -
Mitigating Content Effects on Reasoning in Language Models through Fine-Grained Activation Steering
by: Valentino, Marco, et al.
Published: (2025) -
A Systematic Comparison of Syllogistic Reasoning in Humans and Language Models
by: Eisape, Tiwalayo, et al.
Published: (2023) -
Reasoning with Natural Language Explanations
by: Valentino, Marco, et al.
Published: (2024) -
Binary Autoencoder for Mechanistic Interpretability of Large Language Models
by: Cho, Hakaze, et al.
Published: (2025)