RoFL: Robust Fingerprinting of Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tsai, Yun-Yun, Guo, Chuan, Yang, Junfeng, van der Maaten, Laurens
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909615205122048
author Tsai, Yun-Yun
Guo, Chuan
Yang, Junfeng
van der Maaten, Laurens
author_facet Tsai, Yun-Yun
Guo, Chuan
Yang, Junfeng
van der Maaten, Laurens
contents AI developers are releasing large language models (LLMs) under a variety of different licenses. Many of these licenses restrict the ways in which the models or their outputs may be used. This raises the question how license violations may be recognized. In particular, how can we identify that an API or product uses (an adapted version of) a particular LLM? We present a new method that enable model developers to perform such identification via fingerprints: statistical patterns that are unique to the developer's model and robust to common alterations of that model. Our method permits model identification in a black-box setting using a limited number of queries, enabling identification of models that can only be accessed via an API or product. The fingerprints are non-invasive: our method does not require any changes to the model during training, hence by design, it does not impact model quality. Empirically, we find our method provides a high degree of robustness to common changes in the model or inference settings. In our experiments, it substantially outperforms prior art, including invasive methods that explicitly train watermarks into the model.
format Preprint
id arxiv_https___arxiv_org_abs_2505_12682
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RoFL: Robust Fingerprinting of Language Models
Tsai, Yun-Yun
Guo, Chuan
Yang, Junfeng
van der Maaten, Laurens
Machine Learning
AI developers are releasing large language models (LLMs) under a variety of different licenses. Many of these licenses restrict the ways in which the models or their outputs may be used. This raises the question how license violations may be recognized. In particular, how can we identify that an API or product uses (an adapted version of) a particular LLM? We present a new method that enable model developers to perform such identification via fingerprints: statistical patterns that are unique to the developer's model and robust to common alterations of that model. Our method permits model identification in a black-box setting using a limited number of queries, enabling identification of models that can only be accessed via an API or product. The fingerprints are non-invasive: our method does not require any changes to the model during training, hence by design, it does not impact model quality. Empirically, we find our method provides a high degree of robustness to common changes in the model or inference settings. In our experiments, it substantially outperforms prior art, including invasive methods that explicitly train watermarks into the model.
title RoFL: Robust Fingerprinting of Language Models
topic Machine Learning
url https://arxiv.org/abs/2505.12682