Prompt Compression for Large Language Models: A Survey

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Zongqian, Liu, Yinhong, Su, Yixuan, Collier, Nigel
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912074983014400
author Li, Zongqian
Liu, Yinhong
Su, Yixuan
Collier, Nigel
author_facet Li, Zongqian
Liu, Yinhong
Su, Yixuan
Collier, Nigel
contents Leveraging large language models (LLMs) for complex natural language tasks typically requires long-form prompts to convey detailed requirements and information, which results in increased memory usage and inference costs. To mitigate these challenges, multiple efficient methods have been proposed, with prompt compression gaining significant research interest. This survey provides an overview of prompt compression techniques, categorized into hard prompt methods and soft prompt methods. First, the technical approaches of these methods are compared, followed by an exploration of various ways to understand their mechanisms, including the perspectives of attention optimization, Parameter-Efficient Fine-Tuning (PEFT), modality integration, and new synthetic language. We also examine the downstream adaptations of various prompt compression techniques. Finally, the limitations of current prompt compression methods are analyzed, and several future directions are outlined, such as optimizing the compression encoder, combining hard and soft prompts methods, and leveraging insights from multimodality.
format Preprint
id arxiv_https___arxiv_org_abs_2410_12388
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Prompt Compression for Large Language Models: A Survey
Li, Zongqian
Liu, Yinhong
Su, Yixuan
Collier, Nigel
Computation and Language
Leveraging large language models (LLMs) for complex natural language tasks typically requires long-form prompts to convey detailed requirements and information, which results in increased memory usage and inference costs. To mitigate these challenges, multiple efficient methods have been proposed, with prompt compression gaining significant research interest. This survey provides an overview of prompt compression techniques, categorized into hard prompt methods and soft prompt methods. First, the technical approaches of these methods are compared, followed by an exploration of various ways to understand their mechanisms, including the perspectives of attention optimization, Parameter-Efficient Fine-Tuning (PEFT), modality integration, and new synthetic language. We also examine the downstream adaptations of various prompt compression techniques. Finally, the limitations of current prompt compression methods are analyzed, and several future directions are outlined, such as optimizing the compression encoder, combining hard and soft prompts methods, and leveraging insights from multimodality.
title Prompt Compression for Large Language Models: A Survey
topic Computation and Language
url https://arxiv.org/abs/2410.12388