Fine-Tuning, Quantization, and LLMs: Navigating Unintended Outcomes

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kumar, Divyanshu, Kumar, Anurakt, Agarwal, Sahil, Harshangi, Prashanth
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909308253372416
author Kumar, Divyanshu
Kumar, Anurakt
Agarwal, Sahil
Harshangi, Prashanth
author_facet Kumar, Divyanshu
Kumar, Anurakt
Agarwal, Sahil
Harshangi, Prashanth
contents Large Language Models (LLMs) have gained widespread adoption across various domains, including chatbots and auto-task completion agents. However, these models are susceptible to safety vulnerabilities such as jailbreaking, prompt injection, and privacy leakage attacks. These vulnerabilities can lead to the generation of malicious content, unauthorized actions, or the disclosure of confidential information. While foundational LLMs undergo alignment training and incorporate safety measures, they are often subject to fine-tuning, or doing quantization resource-constrained environments. This study investigates the impact of these modifications on LLM safety, a critical consideration for building reliable and secure AI systems. We evaluate foundational models including Mistral, Llama series, Qwen, and MosaicML, along with their fine-tuned variants. Our comprehensive analysis reveals that fine-tuning generally increases the success rates of jailbreak attacks, while quantization has variable effects on attack success rates. Importantly, we find that properly implemented guardrails significantly enhance resistance to jailbreak attempts. These findings contribute to our understanding of LLM vulnerabilities and provide insights for developing more robust safety strategies in the deployment of language models.
format Preprint
id arxiv_https___arxiv_org_abs_2404_04392
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Fine-Tuning, Quantization, and LLMs: Navigating Unintended Outcomes
Kumar, Divyanshu
Kumar, Anurakt
Agarwal, Sahil
Harshangi, Prashanth
Cryptography and Security
Artificial Intelligence
Large Language Models (LLMs) have gained widespread adoption across various domains, including chatbots and auto-task completion agents. However, these models are susceptible to safety vulnerabilities such as jailbreaking, prompt injection, and privacy leakage attacks. These vulnerabilities can lead to the generation of malicious content, unauthorized actions, or the disclosure of confidential information. While foundational LLMs undergo alignment training and incorporate safety measures, they are often subject to fine-tuning, or doing quantization resource-constrained environments. This study investigates the impact of these modifications on LLM safety, a critical consideration for building reliable and secure AI systems. We evaluate foundational models including Mistral, Llama series, Qwen, and MosaicML, along with their fine-tuned variants. Our comprehensive analysis reveals that fine-tuning generally increases the success rates of jailbreak attacks, while quantization has variable effects on attack success rates. Importantly, we find that properly implemented guardrails significantly enhance resistance to jailbreak attempts. These findings contribute to our understanding of LLM vulnerabilities and provide insights for developing more robust safety strategies in the deployment of language models.
title Fine-Tuning, Quantization, and LLMs: Navigating Unintended Outcomes
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2404.04392