MEraser: An Effective Fingerprint Erasure Approach for Large Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhang, Jingxuan, Xu, Zhenhua, Hu, Rui, Xing, Wenpeng, Zhang, Xuhong, Han, Meng
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916920162254848
author Zhang, Jingxuan
Xu, Zhenhua
Hu, Rui
Xing, Wenpeng
Zhang, Xuhong
Han, Meng
author_facet Zhang, Jingxuan
Xu, Zhenhua
Hu, Rui
Xing, Wenpeng
Zhang, Xuhong
Han, Meng
contents Large Language Models (LLMs) have become increasingly prevalent across various sectors, raising critical concerns about model ownership and intellectual property protection. Although backdoor-based fingerprinting has emerged as a promising solution for model authentication, effective attacks for removing these fingerprints remain largely unexplored. Therefore, we present Mismatched Eraser (MEraser), a novel method for effectively removing backdoor-based fingerprints from LLMs while maintaining model performance. Our approach leverages a two-phase fine-tuning strategy utilizing carefully constructed mismatched and clean datasets. Through extensive evaluation across multiple LLM architectures and fingerprinting methods, we demonstrate that MEraser achieves complete fingerprinting removal while maintaining model performance with minimal training data of fewer than 1,000 samples. Furthermore, we introduce a transferable erasure mechanism that enables effective fingerprinting removal across different models without repeated training. In conclusion, our approach provides a practical solution for fingerprinting removal in LLMs, reveals critical vulnerabilities in current fingerprinting techniques, and establishes comprehensive evaluation benchmarks for developing more resilient model protection methods in the future.
format Preprint
id arxiv_https___arxiv_org_abs_2506_12551
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MEraser: An Effective Fingerprint Erasure Approach for Large Language Models
Zhang, Jingxuan
Xu, Zhenhua
Hu, Rui
Xing, Wenpeng
Zhang, Xuhong
Han, Meng
Cryptography and Security
Artificial Intelligence
Large Language Models (LLMs) have become increasingly prevalent across various sectors, raising critical concerns about model ownership and intellectual property protection. Although backdoor-based fingerprinting has emerged as a promising solution for model authentication, effective attacks for removing these fingerprints remain largely unexplored. Therefore, we present Mismatched Eraser (MEraser), a novel method for effectively removing backdoor-based fingerprints from LLMs while maintaining model performance. Our approach leverages a two-phase fine-tuning strategy utilizing carefully constructed mismatched and clean datasets. Through extensive evaluation across multiple LLM architectures and fingerprinting methods, we demonstrate that MEraser achieves complete fingerprinting removal while maintaining model performance with minimal training data of fewer than 1,000 samples. Furthermore, we introduce a transferable erasure mechanism that enables effective fingerprinting removal across different models without repeated training. In conclusion, our approach provides a practical solution for fingerprinting removal in LLMs, reveals critical vulnerabilities in current fingerprinting techniques, and establishes comprehensive evaluation benchmarks for developing more resilient model protection methods in the future.
title MEraser: An Effective Fingerprint Erasure Approach for Large Language Models
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2506.12551