AttnDiff: Attention-based Differential Fingerprinting for Large Language Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhang, Haobo, Xu, Zhenhua, Li, Junxian, Sheng, Shangfeng, Kong, Dezhang, Han, Meng
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913009615503360
author Zhang, Haobo
Xu, Zhenhua
Li, Junxian
Sheng, Shangfeng
Kong, Dezhang
Han, Meng
author_facet Zhang, Haobo
Xu, Zhenhua
Li, Junxian
Sheng, Shangfeng
Kong, Dezhang
Han, Meng
contents Protecting the intellectual property of open-weight large language models (LLMs) requires verifying whether a suspect model is derived from a victim model despite common laundering operations such as fine-tuning (including PPO/DPO), pruning/compression, and model merging. We propose \textsc{AttnDiff}, a data-efficient white-box framework that extracts fingerprints from models via intrinsic information-routing behavior. \textsc{AttnDiff} probes minimally edited prompt pairs that induce controlled semantic conflicts, captures differential attention patterns, summarizes them with compact spectral descriptors, and compares models using CKA. Across Llama-2/3 and Qwen2.5 (3B--14B) and additional open-source families, it yields high similarity for related derivatives while separating unrelated model families (e.g., $>0.98$ vs.\ $<0.22$ with $M=60$ probes). With 5--60 multi-domain probes, it supports practical provenance verification and accountability.
format Preprint
id arxiv_https___arxiv_org_abs_2604_05502
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AttnDiff: Attention-based Differential Fingerprinting for Large Language Models
Zhang, Haobo
Xu, Zhenhua
Li, Junxian
Sheng, Shangfeng
Kong, Dezhang
Han, Meng
Cryptography and Security
Machine Learning
Protecting the intellectual property of open-weight large language models (LLMs) requires verifying whether a suspect model is derived from a victim model despite common laundering operations such as fine-tuning (including PPO/DPO), pruning/compression, and model merging. We propose \textsc{AttnDiff}, a data-efficient white-box framework that extracts fingerprints from models via intrinsic information-routing behavior. \textsc{AttnDiff} probes minimally edited prompt pairs that induce controlled semantic conflicts, captures differential attention patterns, summarizes them with compact spectral descriptors, and compares models using CKA. Across Llama-2/3 and Qwen2.5 (3B--14B) and additional open-source families, it yields high similarity for related derivatives while separating unrelated model families (e.g., $>0.98$ vs.\ $<0.22$ with $M=60$ probes). With 5--60 multi-domain probes, it supports practical provenance verification and accountability.
title AttnDiff: Attention-based Differential Fingerprinting for Large Language Models
topic Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2604.05502