LoRAShield: Data-Free Editing Alignment for Secure Personalized LoRA Sharing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Jiahao, li, junhao, Wang, Yiming, Ma, Zhe, Jiang, Yi, Zhou, Chunyi, Li, Qingming, Du, Tianyu, Ji, Shouling
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909681591517184
author Chen, Jiahao
li, junhao
Wang, Yiming
Ma, Zhe
Jiang, Yi
Zhou, Chunyi
Li, Qingming
Du, Tianyu
Ji, Shouling
author_facet Chen, Jiahao
li, junhao
Wang, Yiming
Ma, Zhe
Jiang, Yi
Zhou, Chunyi
Li, Qingming
Du, Tianyu
Ji, Shouling
contents The proliferation of Low-Rank Adaptation (LoRA) models has democratized personalized text-to-image generation, enabling users to share lightweight models (e.g., personal portraits) on platforms like Civitai and Liblib. However, this "share-and-play" ecosystem introduces critical risks: benign LoRAs can be weaponized by adversaries to generate harmful content (e.g., political, defamatory imagery), undermining creator rights and platform safety. Existing defenses like concept-erasure methods focus on full diffusion models (DMs), neglecting LoRA's unique role as a modular adapter and its vulnerability to adversarial prompt engineering. To bridge this gap, we propose LoRAShield, the first data-free editing framework for securing LoRA models against misuse. Our platform-driven approach dynamically edits and realigns LoRA's weight subspace via adversarial optimization and semantic augmentation. Experimental results demonstrate that LoRAShield achieves remarkable effectiveness, efficiency, and robustness in blocking malicious generations without sacrificing the functionality of the benign task. By shifting the defense to platforms, LoRAShield enables secure, scalable sharing of personalized models, a critical step toward trustworthy generative ecosystems.
format Preprint
id arxiv_https___arxiv_org_abs_2507_07056
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LoRAShield: Data-Free Editing Alignment for Secure Personalized LoRA Sharing
Chen, Jiahao
li, junhao
Wang, Yiming
Ma, Zhe
Jiang, Yi
Zhou, Chunyi
Li, Qingming
Du, Tianyu
Ji, Shouling
Cryptography and Security
Machine Learning
The proliferation of Low-Rank Adaptation (LoRA) models has democratized personalized text-to-image generation, enabling users to share lightweight models (e.g., personal portraits) on platforms like Civitai and Liblib. However, this "share-and-play" ecosystem introduces critical risks: benign LoRAs can be weaponized by adversaries to generate harmful content (e.g., political, defamatory imagery), undermining creator rights and platform safety. Existing defenses like concept-erasure methods focus on full diffusion models (DMs), neglecting LoRA's unique role as a modular adapter and its vulnerability to adversarial prompt engineering. To bridge this gap, we propose LoRAShield, the first data-free editing framework for securing LoRA models against misuse. Our platform-driven approach dynamically edits and realigns LoRA's weight subspace via adversarial optimization and semantic augmentation. Experimental results demonstrate that LoRAShield achieves remarkable effectiveness, efficiency, and robustness in blocking malicious generations without sacrificing the functionality of the benign task. By shifting the defense to platforms, LoRAShield enables secure, scalable sharing of personalized models, a critical step toward trustworthy generative ecosystems.
title LoRAShield: Data-Free Editing Alignment for Secure Personalized LoRA Sharing
topic Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2507.07056