REEF: Representation Encoding Fingerprints for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Jie, Liu, Dongrui, Qian, Chen, Zhang, Linfeng, Liu, Yong, Qiao, Yu, Shao, Jing
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913553220370432
author Zhang, Jie
Liu, Dongrui
Qian, Chen
Zhang, Linfeng
Liu, Yong
Qiao, Yu
Shao, Jing
author_facet Zhang, Jie
Liu, Dongrui
Qian, Chen
Zhang, Linfeng
Liu, Yong
Qiao, Yu
Shao, Jing
contents Protecting the intellectual property of open-source Large Language Models (LLMs) is very important, because training LLMs costs extensive computational resources and data. Therefore, model owners and third parties need to identify whether a suspect model is a subsequent development of the victim model. To this end, we propose a training-free REEF to identify the relationship between the suspect and victim models from the perspective of LLMs' feature representations. Specifically, REEF computes and compares the centered kernel alignment similarity between the representations of a suspect model and a victim model on the same samples. This training-free REEF does not impair the model's general capabilities and is robust to sequential fine-tuning, pruning, model merging, and permutations. In this way, REEF provides a simple and effective way for third parties and models' owners to protect LLMs' intellectual property together. The code is available at https://github.com/tmylla/REEF.
format Preprint
id arxiv_https___arxiv_org_abs_2410_14273
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle REEF: Representation Encoding Fingerprints for Large Language Models
Zhang, Jie
Liu, Dongrui
Qian, Chen
Zhang, Linfeng
Liu, Yong
Qiao, Yu
Shao, Jing
Computation and Language
Artificial Intelligence
Cryptography and Security
Protecting the intellectual property of open-source Large Language Models (LLMs) is very important, because training LLMs costs extensive computational resources and data. Therefore, model owners and third parties need to identify whether a suspect model is a subsequent development of the victim model. To this end, we propose a training-free REEF to identify the relationship between the suspect and victim models from the perspective of LLMs' feature representations. Specifically, REEF computes and compares the centered kernel alignment similarity between the representations of a suspect model and a victim model on the same samples. This training-free REEF does not impair the model's general capabilities and is robust to sequential fine-tuning, pruning, model merging, and permutations. In this way, REEF provides a simple and effective way for third parties and models' owners to protect LLMs' intellectual property together. The code is available at https://github.com/tmylla/REEF.
title REEF: Representation Encoding Fingerprints for Large Language Models
topic Computation and Language
Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2410.14273