Anchor Attention, Small Cache: Code Generation with Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Xiangyu, Zhou, Yu, Yang, Guang, Gall, Harald C., Chen, Taolue
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929586918391808
author Zhang, Xiangyu
Zhou, Yu
Yang, Guang
Gall, Harald C.
Chen, Taolue
author_facet Zhang, Xiangyu
Zhou, Yu
Yang, Guang
Gall, Harald C.
Chen, Taolue
contents The development of large language models (LLMs) has revolutionized automated code generation. However, their high demand of computation resources has hindered a broader deployment and raised environmental concerns. A common strategy for diminishing computational demands is to cache Key-Value (KV) states from the attention mechanism which is adopted predominately by mainstream LLMs. It can mitigate the need of repeated attention computations, but brings significant memory overhead. Current practices in NLP often use sparse attention which may, unfortunately, lead to substantial inaccuracies, or hallucinations, in code generation tasks. In this paper, we analyze the attention weights distribution within code generation models via an empirical study, uncovering a sparsity pattern, i.e., the aggregation of information at specific anchor points. Based on this observation, we propose a novel approach, AnchorCoder, which features token-wise anchor attention designed to extract and compress the contextual information, and layer-wise anchor attention enabling cross-layer communication to mitigate the issue of excessive superposition caused by the compression. The extensive experiments across multiple benchmark datasets confirm the effectiveness of AnchorCoder, which can consistently achieve a significant (at least 70%) reduction in KV cache requirements, while preserving the majority of model's performance.
format Preprint
id arxiv_https___arxiv_org_abs_2411_06680
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Anchor Attention, Small Cache: Code Generation with Large Language Models
Zhang, Xiangyu
Zhou, Yu
Yang, Guang
Gall, Harald C.
Chen, Taolue
Software Engineering
68N19
D.2.3
The development of large language models (LLMs) has revolutionized automated code generation. However, their high demand of computation resources has hindered a broader deployment and raised environmental concerns. A common strategy for diminishing computational demands is to cache Key-Value (KV) states from the attention mechanism which is adopted predominately by mainstream LLMs. It can mitigate the need of repeated attention computations, but brings significant memory overhead. Current practices in NLP often use sparse attention which may, unfortunately, lead to substantial inaccuracies, or hallucinations, in code generation tasks. In this paper, we analyze the attention weights distribution within code generation models via an empirical study, uncovering a sparsity pattern, i.e., the aggregation of information at specific anchor points. Based on this observation, we propose a novel approach, AnchorCoder, which features token-wise anchor attention designed to extract and compress the contextual information, and layer-wise anchor attention enabling cross-layer communication to mitigate the issue of excessive superposition caused by the compression. The extensive experiments across multiple benchmark datasets confirm the effectiveness of AnchorCoder, which can consistently achieve a significant (at least 70%) reduction in KV cache requirements, while preserving the majority of model's performance.
title Anchor Attention, Small Cache: Code Generation with Large Language Models
topic Software Engineering
68N19
D.2.3
url https://arxiv.org/abs/2411.06680