UniAttn: Reducing Inference Costs via Softmax Unification for Post-Training LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917216173162496 |
|---|---|
| author | Xiong, Yizhe Huang, Wei Ye, Xin Chen, Hui Lin, Zijia Lian, Haoran Su, Zhenpeng Han, Jungong Ding, Guiguang |
| author_facet | Xiong, Yizhe Huang, Wei Ye, Xin Chen, Hui Lin, Zijia Lian, Haoran Su, Zhenpeng Han, Jungong Ding, Guiguang |
| contents | Post-training is essential for adapting Large Language Models (LLMs) to real-world applications. Deploying post-trained models faces significant challenges due to substantial memory overhead and noticeable inference latency. Existing work has identified significant redundancies in LLMs and proposed efficient architectures, namely intra-layer KV sharing and cross-layer KV sharing. However, these methods still result in high inference time overhead, remaining suboptimal for post-training pre-trained LLMs. In this paper, we identify that the \texttt{Softmax} operation is a primary bottleneck for LLM inference and discover that it is actually highly redundant during post-training. We propose Softmax \textbf{Uni}fication in \textbf{Att}e\textbf{n}tion (\textbf{UniAttn}), a novel post-training method that unifies Softmax activations across transformer blocks to reduce LLM inference costs. Additionally, UniAttn adopts a linear projection to compensate for the errors induced by Softmax unification. Experiments show that UniAttn matches the performance of standard post-training while significantly reducing inference costs, outperforming existing efficient architectures during post-training. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2502_00439 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | UniAttn: Reducing Inference Costs via Softmax Unification for Post-Training LLMs Xiong, Yizhe Huang, Wei Ye, Xin Chen, Hui Lin, Zijia Lian, Haoran Su, Zhenpeng Han, Jungong Ding, Guiguang Computation and Language Post-training is essential for adapting Large Language Models (LLMs) to real-world applications. Deploying post-trained models faces significant challenges due to substantial memory overhead and noticeable inference latency. Existing work has identified significant redundancies in LLMs and proposed efficient architectures, namely intra-layer KV sharing and cross-layer KV sharing. However, these methods still result in high inference time overhead, remaining suboptimal for post-training pre-trained LLMs. In this paper, we identify that the \texttt{Softmax} operation is a primary bottleneck for LLM inference and discover that it is actually highly redundant during post-training. We propose Softmax \textbf{Uni}fication in \textbf{Att}e\textbf{n}tion (\textbf{UniAttn}), a novel post-training method that unifies Softmax activations across transformer blocks to reduce LLM inference costs. Additionally, UniAttn adopts a linear projection to compensate for the errors induced by Softmax unification. Experiments show that UniAttn matches the performance of standard post-training while significantly reducing inference costs, outperforming existing efficient architectures during post-training. |
| title | UniAttn: Reducing Inference Costs via Softmax Unification for Post-Training LLMs |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2502.00439 |