Layer of Truth: Probing Belief Shifts under Continual Pre-Training Poisoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Churina, Svetlana, Chebrolu, Niranjan, Jaidka, Kokil
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917253567479808
author Churina, Svetlana
Chebrolu, Niranjan
Jaidka, Kokil
author_facet Churina, Svetlana
Chebrolu, Niranjan
Jaidka, Kokil
contents We show that continual pretraining on plausible misinformation can overwrite specific factual knowledge in large language models without degrading overall performance. Unlike prior poisoning work under static pretraining, we study repeated exposure to counterfactual claims during continual updates. Using paired fact-counterfact items with graded poisoning ratios, we track how internal preferences between competing facts evolve across checkpoints, layers, and model scales. Even moderate poisoning (50-100%) flips over 55% of responses from correct to counterfactual while leaving ambiguity nearly unchanged. These belief flips emerge abruptly, concentrate in late layers (e.g., Layers 29-36 in 3B models), and are partially reversible via patching (up to 56.8%). The corrupted beliefs generalize beyond poisoned prompts, selectively degrading commonsense reasoning while leaving alignment benchmarks largely intact and transferring imperfectly across languages. These results expose a failure mode of continual pre-training in which targeted misinformation replaces internal factual representations without triggering broad performance collapse, motivating representation-level monitoring of factual integrity during model updates.
format Preprint
id arxiv_https___arxiv_org_abs_2510_26829
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Layer of Truth: Probing Belief Shifts under Continual Pre-Training Poisoning
Churina, Svetlana
Chebrolu, Niranjan
Jaidka, Kokil
Machine Learning
Cryptography and Security
We show that continual pretraining on plausible misinformation can overwrite specific factual knowledge in large language models without degrading overall performance. Unlike prior poisoning work under static pretraining, we study repeated exposure to counterfactual claims during continual updates. Using paired fact-counterfact items with graded poisoning ratios, we track how internal preferences between competing facts evolve across checkpoints, layers, and model scales. Even moderate poisoning (50-100%) flips over 55% of responses from correct to counterfactual while leaving ambiguity nearly unchanged. These belief flips emerge abruptly, concentrate in late layers (e.g., Layers 29-36 in 3B models), and are partially reversible via patching (up to 56.8%). The corrupted beliefs generalize beyond poisoned prompts, selectively degrading commonsense reasoning while leaving alignment benchmarks largely intact and transferring imperfectly across languages. These results expose a failure mode of continual pre-training in which targeted misinformation replaces internal factual representations without triggering broad performance collapse, motivating representation-level monitoring of factual integrity during model updates.
title Layer of Truth: Probing Belief Shifts under Continual Pre-Training Poisoning
topic Machine Learning
Cryptography and Security
url https://arxiv.org/abs/2510.26829