UniAttn: Reducing Inference Costs via Softmax Unification for Post-Training LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xiong, Yizhe, Huang, Wei, Ye, Xin, Chen, Hui, Lin, Zijia, Lian, Haoran, Su, Zhenpeng, Han, Jungong, Ding, Guiguang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917216173162496
author Xiong, Yizhe
Huang, Wei
Ye, Xin
Chen, Hui
Lin, Zijia
Lian, Haoran
Su, Zhenpeng
Han, Jungong
Ding, Guiguang
author_facet Xiong, Yizhe
Huang, Wei
Ye, Xin
Chen, Hui
Lin, Zijia
Lian, Haoran
Su, Zhenpeng
Han, Jungong
Ding, Guiguang
contents Post-training is essential for adapting Large Language Models (LLMs) to real-world applications. Deploying post-trained models faces significant challenges due to substantial memory overhead and noticeable inference latency. Existing work has identified significant redundancies in LLMs and proposed efficient architectures, namely intra-layer KV sharing and cross-layer KV sharing. However, these methods still result in high inference time overhead, remaining suboptimal for post-training pre-trained LLMs. In this paper, we identify that the \texttt{Softmax} operation is a primary bottleneck for LLM inference and discover that it is actually highly redundant during post-training. We propose Softmax \textbf{Uni}fication in \textbf{Att}e\textbf{n}tion (\textbf{UniAttn}), a novel post-training method that unifies Softmax activations across transformer blocks to reduce LLM inference costs. Additionally, UniAttn adopts a linear projection to compensate for the errors induced by Softmax unification. Experiments show that UniAttn matches the performance of standard post-training while significantly reducing inference costs, outperforming existing efficient architectures during post-training.
format Preprint
id arxiv_https___arxiv_org_abs_2502_00439
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle UniAttn: Reducing Inference Costs via Softmax Unification for Post-Training LLMs
Xiong, Yizhe
Huang, Wei
Ye, Xin
Chen, Hui
Lin, Zijia
Lian, Haoran
Su, Zhenpeng
Han, Jungong
Ding, Guiguang
Computation and Language
Post-training is essential for adapting Large Language Models (LLMs) to real-world applications. Deploying post-trained models faces significant challenges due to substantial memory overhead and noticeable inference latency. Existing work has identified significant redundancies in LLMs and proposed efficient architectures, namely intra-layer KV sharing and cross-layer KV sharing. However, these methods still result in high inference time overhead, remaining suboptimal for post-training pre-trained LLMs. In this paper, we identify that the \texttt{Softmax} operation is a primary bottleneck for LLM inference and discover that it is actually highly redundant during post-training. We propose Softmax \textbf{Uni}fication in \textbf{Att}e\textbf{n}tion (\textbf{UniAttn}), a novel post-training method that unifies Softmax activations across transformer blocks to reduce LLM inference costs. Additionally, UniAttn adopts a linear projection to compensate for the errors induced by Softmax unification. Experiments show that UniAttn matches the performance of standard post-training while significantly reducing inference costs, outperforming existing efficient architectures during post-training.
title UniAttn: Reducing Inference Costs via Softmax Unification for Post-Training LLMs
topic Computation and Language
url https://arxiv.org/abs/2502.00439