Fingerprinting LLMs via Prompt Injection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Yuepeng, Jiang, Zhengyuan, Li, Mengyuan, Ahmed, Osama, Huang, Zhicong, Hong, Cheng, Gong, Neil
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910233374228480
author Hu, Yuepeng
Jiang, Zhengyuan
Li, Mengyuan
Ahmed, Osama
Huang, Zhicong
Hong, Cheng
Gong, Neil
author_facet Hu, Yuepeng
Jiang, Zhengyuan
Li, Mengyuan
Ahmed, Osama
Huang, Zhicong
Hong, Cheng
Gong, Neil
contents Large language models (LLMs) are often modified after release through post-processing such as post-training or quantization, which makes it challenging to determine whether one model is derived from another. Existing provenance detection methods have two main limitations: (1) they embed signals into the base model before release, which is infeasible for already published models, or (2) they compare outputs across models using hand-crafted or random prompts, which are not robust to post-processing. In this work, we propose LLMPrint, a novel detection framework that constructs fingerprints by exploiting LLMs' inherent vulnerability to prompt injection. Our key insight is that by optimizing fingerprint prompts to enforce consistent token preferences, we can obtain fingerprints that are both unique to the base model and robust to post-processing. We further develop a unified verification procedure that applies to both gray-box and black-box settings, with statistical guarantees. We evaluate LLMPrint on five base models and around 700 post-trained or quantized variants. Our results show that LLMPrint achieves high true positive rates while keeping false positive rates near zero. The code is publicly available at https://github.com/hifi-hyp/ACL-LLMPrint.
format Preprint
id arxiv_https___arxiv_org_abs_2509_25448
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Fingerprinting LLMs via Prompt Injection
Hu, Yuepeng
Jiang, Zhengyuan
Li, Mengyuan
Ahmed, Osama
Huang, Zhicong
Hong, Cheng
Gong, Neil
Cryptography and Security
Computation and Language
Large language models (LLMs) are often modified after release through post-processing such as post-training or quantization, which makes it challenging to determine whether one model is derived from another. Existing provenance detection methods have two main limitations: (1) they embed signals into the base model before release, which is infeasible for already published models, or (2) they compare outputs across models using hand-crafted or random prompts, which are not robust to post-processing. In this work, we propose LLMPrint, a novel detection framework that constructs fingerprints by exploiting LLMs' inherent vulnerability to prompt injection. Our key insight is that by optimizing fingerprint prompts to enforce consistent token preferences, we can obtain fingerprints that are both unique to the base model and robust to post-processing. We further develop a unified verification procedure that applies to both gray-box and black-box settings, with statistical guarantees. We evaluate LLMPrint on five base models and around 700 post-trained or quantized variants. Our results show that LLMPrint achieves high true positive rates while keeping false positive rates near zero. The code is publicly available at https://github.com/hifi-hyp/ACL-LLMPrint.
title Fingerprinting LLMs via Prompt Injection
topic Cryptography and Security
Computation and Language
url https://arxiv.org/abs/2509.25448