LeftoverLocals: Listening to LLM Responses Through Leaked GPU Local Memory

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sorensen, Tyler, Khlaaf, Heidy
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911767466082304
author Sorensen, Tyler
Khlaaf, Heidy
author_facet Sorensen, Tyler
Khlaaf, Heidy
contents This paper describes LeftoverLocals: a vulnerability that allows data recovery from GPU memory created by another process on Apple, Qualcomm, and AMD GPUs. LeftoverLocals impacts the security posture of GPU applications, with particular significance to LLMs and ML models that run on impacted GPUs. By recovering local memory, an optimized GPU memory region, we built a PoC where an attacker can listen into another user's interactive LLM session (e.g., llama.cpp) across process or container boundaries.
format Preprint
id arxiv_https___arxiv_org_abs_2401_16603
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LeftoverLocals: Listening to LLM Responses Through Leaked GPU Local Memory
Sorensen, Tyler
Khlaaf, Heidy
Cryptography and Security
Distributed, Parallel, and Cluster Computing
This paper describes LeftoverLocals: a vulnerability that allows data recovery from GPU memory created by another process on Apple, Qualcomm, and AMD GPUs. LeftoverLocals impacts the security posture of GPU applications, with particular significance to LLMs and ML models that run on impacted GPUs. By recovering local memory, an optimized GPU memory region, we built a PoC where an attacker can listen into another user's interactive LLM session (e.g., llama.cpp) across process or container boundaries.
title LeftoverLocals: Listening to LLM Responses Through Leaked GPU Local Memory
topic Cryptography and Security
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2401.16603