Tracing Facts or just Copies? A critical investigation of the Competitions of Mechanisms in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Campregher, Dante, Chen, Yanxu, Hoffman, Sander, Heuss, Maria
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911059170820096
author Campregher, Dante
Chen, Yanxu
Hoffman, Sander
Heuss, Maria
author_facet Campregher, Dante
Chen, Yanxu
Hoffman, Sander
Heuss, Maria
contents This paper presents a reproducibility study examining how Large Language Models (LLMs) manage competing factual and counterfactual information, focusing on the role of attention heads in this process. We attempt to reproduce and reconcile findings from three recent studies by Ortu et al., Yu, Merullo, and Pavlick and McDougall et al. that investigate the competition between model-learned facts and contradictory context information through Mechanistic Interpretability tools. Our study specifically examines the relationship between attention head strength and factual output ratios, evaluates competing hypotheses about attention heads' suppression mechanisms, and investigates the domain specificity of these attention patterns. Our findings suggest that attention heads promoting factual output do so via general copy suppression rather than selective counterfactual suppression, as strengthening them can also inhibit correct facts. Additionally, we show that attention head behavior is domain-dependent, with larger models exhibiting more specialized and category-sensitive patterns.
format Preprint
id arxiv_https___arxiv_org_abs_2507_11809
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Tracing Facts or just Copies? A critical investigation of the Competitions of Mechanisms in Large Language Models
Campregher, Dante
Chen, Yanxu
Hoffman, Sander
Heuss, Maria
Computation and Language
Artificial Intelligence
Machine Learning
This paper presents a reproducibility study examining how Large Language Models (LLMs) manage competing factual and counterfactual information, focusing on the role of attention heads in this process. We attempt to reproduce and reconcile findings from three recent studies by Ortu et al., Yu, Merullo, and Pavlick and McDougall et al. that investigate the competition between model-learned facts and contradictory context information through Mechanistic Interpretability tools. Our study specifically examines the relationship between attention head strength and factual output ratios, evaluates competing hypotheses about attention heads' suppression mechanisms, and investigates the domain specificity of these attention patterns. Our findings suggest that attention heads promoting factual output do so via general copy suppression rather than selective counterfactual suppression, as strengthening them can also inhibit correct facts. Additionally, we show that attention head behavior is domain-dependent, with larger models exhibiting more specialized and category-sensitive patterns.
title Tracing Facts or just Copies? A critical investigation of the Competitions of Mechanisms in Large Language Models
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2507.11809