Leaking LoRa: An Evaluation of Password Leaks and Knowledge Storage in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Marinelli, Ryan, Eckhoff, Magnus
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908292381409280
author Marinelli, Ryan
Eckhoff, Magnus
author_facet Marinelli, Ryan
Eckhoff, Magnus
contents To effectively deploy Large Language Models (LLMs) in application-specific settings, fine-tuning techniques are applied to enhance performance on specialized tasks. This process often involves fine-tuning on user data data, which may contain sensitive information. Although not recommended, it is not uncommon for users to send passwords in messages, and fine-tuning models on this could result in passwords being leaked. In this study, a Large Language Model is fine-tuned with customer support data and passwords from the RockYou password wordlist using Low-Rank Adaptation (LoRA). Out of the first 200 passwords from the list, 37 were successfully recovered. Further, causal tracing is used to identify that password information is largely located in a few layers. Lastly, Rank One Model Editing (ROME) is used to remove the password information from the model, resulting in the number of passwords recovered going from 37 to 0.
format Preprint
id arxiv_https___arxiv_org_abs_2504_00031
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Leaking LoRa: An Evaluation of Password Leaks and Knowledge Storage in Large Language Models
Marinelli, Ryan
Eckhoff, Magnus
Cryptography and Security
Artificial Intelligence
Computation and Language
To effectively deploy Large Language Models (LLMs) in application-specific settings, fine-tuning techniques are applied to enhance performance on specialized tasks. This process often involves fine-tuning on user data data, which may contain sensitive information. Although not recommended, it is not uncommon for users to send passwords in messages, and fine-tuning models on this could result in passwords being leaked. In this study, a Large Language Model is fine-tuned with customer support data and passwords from the RockYou password wordlist using Low-Rank Adaptation (LoRA). Out of the first 200 passwords from the list, 37 were successfully recovered. Further, causal tracing is used to identify that password information is largely located in a few layers. Lastly, Rank One Model Editing (ROME) is used to remove the password information from the model, resulting in the number of passwords recovered going from 37 to 0.
title Leaking LoRa: An Evaluation of Password Leaks and Knowledge Storage in Large Language Models
topic Cryptography and Security
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2504.00031