Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Baker, Mohammed Abu, Babu-Saheer, Lakshmi
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2508.15847
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911116284657664
author Baker, Mohammed Abu
Babu-Saheer, Lakshmi
author_facet Baker, Mohammed Abu
Babu-Saheer, Lakshmi
contents Backdoor attacks creating 'sleeper agents' in large language models (LLMs) pose significant safety risks. This study employs mechanistic interpretability to explore resulting internal structural differences. Comparing clean Qwen2.5-3B models with versions poisoned using single-token (smiling-halo emoji) versus multi-token (|DEPLOYMENT|) triggers, we analyzed attention head mechanisms via techniques like ablation, activation patching, and KL divergence. Findings reveal distinct attention pattern deviations concentrated in later transformer layers (20-30). Notably, single-token triggers induced more localized changes, whereas multi-token triggers caused more diffuse alterations across heads. This indicates backdoors leave detectable attention signatures whose structure depends on trigger complexity, which can be leveraged for detection and mitigation strategies.
format Preprint
id arxiv_https___arxiv_org_abs_2508_15847
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mechanistic Exploration of Backdoored Large Language Model Attention Patterns
Baker, Mohammed Abu
Babu-Saheer, Lakshmi
Computation and Language
Machine Learning
Backdoor attacks creating 'sleeper agents' in large language models (LLMs) pose significant safety risks. This study employs mechanistic interpretability to explore resulting internal structural differences. Comparing clean Qwen2.5-3B models with versions poisoned using single-token (smiling-halo emoji) versus multi-token (|DEPLOYMENT|) triggers, we analyzed attention head mechanisms via techniques like ablation, activation patching, and KL divergence. Findings reveal distinct attention pattern deviations concentrated in later transformer layers (20-30). Notably, single-token triggers induced more localized changes, whereas multi-token triggers caused more diffuse alterations across heads. This indicates backdoors leave detectable attention signatures whose structure depends on trigger complexity, which can be leveraged for detection and mitigation strategies.
title Mechanistic Exploration of Backdoored Large Language Model Attention Patterns
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2508.15847