ENEC: A Lossless AI Model Compression Method Enabling Fast Inference on Ascend NPUs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Jinwu, Wu, Jiaan, Liu, Zedong, Ma, Xinyang, Zhao, Hairui, Gu, Yida, Huang, Yuanhong, Liu, Xingchen, Huang, Wenjing, Wei, Zheng, Xing, Jing, Ma, Yili, Zhang, Qingyi, An, Baoyi, Hu, Zhongzhe, Liu, Shaoteng, Zhu, Xia, Lu, Jiaxun, Tan, Guangming, Tao, Dingwen
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911573415559168
author Yang, Jinwu
Wu, Jiaan
Liu, Zedong
Ma, Xinyang
Zhao, Hairui
Gu, Yida
Huang, Yuanhong
Liu, Xingchen
Huang, Wenjing
Wei, Zheng
Xing, Jing
Ma, Yili
Zhang, Qingyi
An, Baoyi
Hu, Zhongzhe
Liu, Shaoteng
Zhu, Xia
Lu, Jiaxun
Tan, Guangming
Tao, Dingwen
author_facet Yang, Jinwu
Wu, Jiaan
Liu, Zedong
Ma, Xinyang
Zhao, Hairui
Gu, Yida
Huang, Yuanhong
Liu, Xingchen
Huang, Wenjing
Wei, Zheng
Xing, Jing
Ma, Yili
Zhang, Qingyi
An, Baoyi
Hu, Zhongzhe
Liu, Shaoteng
Zhu, Xia
Lu, Jiaxun
Tan, Guangming
Tao, Dingwen
contents The rapid scaling of Large Language Models presents significant challenges for their deployment and inference, particularly on resource-constrained specialized AI hardware accelerators such as Huawei's Ascend NPUs, where weight data transfer has become a critical performance bottleneck. While lossless compression can preserve model accuracy and reduce data volume, existing lossless compression algorithms exhibit extremely low throughput when ported to the Ascend NPU architecture. In this paper, we propose ENEC, a novel lossless compression method specifically customized for AI model weights and optimized for Ascend Neural Processing Units. ENEC adopts a block-based fixed-length encoding scheme and incorporates a series of NPU-specific optimizations: bit-width quantization with hierarchical halving bit-packing, vectorized branch-free integer transformation, and dependency-decoupled intra-segment scan for efficient prefix-sum computation. Experimental results demonstrate that ENEC outperforms existing state-of-the-art NPU compressors in both compression ratio and throughput. Compared to leading GPU solutions, ENEC achieves a 3.43X higher throughput than DietGPU and a 1.12X better compression ratio than nvCOMP. By reducing weight transmission overhead, ENEC significantly improves end-to-end inference performance, achieving up to a 6.3X speedup. On Ascend NPUs, ENEC is the first open-source lossless compression algorithm for model weights that achieves performance comparable to state-of-the-art GPU compressors, offering an effective solution for deploying large-scale AI models.
format Preprint
id arxiv_https___arxiv_org_abs_2604_03298
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ENEC: A Lossless AI Model Compression Method Enabling Fast Inference on Ascend NPUs
Yang, Jinwu
Wu, Jiaan
Liu, Zedong
Ma, Xinyang
Zhao, Hairui
Gu, Yida
Huang, Yuanhong
Liu, Xingchen
Huang, Wenjing
Wei, Zheng
Xing, Jing
Ma, Yili
Zhang, Qingyi
An, Baoyi
Hu, Zhongzhe
Liu, Shaoteng
Zhu, Xia
Lu, Jiaxun
Tan, Guangming
Tao, Dingwen
Hardware Architecture
Distributed, Parallel, and Cluster Computing
Machine Learning
The rapid scaling of Large Language Models presents significant challenges for their deployment and inference, particularly on resource-constrained specialized AI hardware accelerators such as Huawei's Ascend NPUs, where weight data transfer has become a critical performance bottleneck. While lossless compression can preserve model accuracy and reduce data volume, existing lossless compression algorithms exhibit extremely low throughput when ported to the Ascend NPU architecture. In this paper, we propose ENEC, a novel lossless compression method specifically customized for AI model weights and optimized for Ascend Neural Processing Units. ENEC adopts a block-based fixed-length encoding scheme and incorporates a series of NPU-specific optimizations: bit-width quantization with hierarchical halving bit-packing, vectorized branch-free integer transformation, and dependency-decoupled intra-segment scan for efficient prefix-sum computation. Experimental results demonstrate that ENEC outperforms existing state-of-the-art NPU compressors in both compression ratio and throughput. Compared to leading GPU solutions, ENEC achieves a 3.43X higher throughput than DietGPU and a 1.12X better compression ratio than nvCOMP. By reducing weight transmission overhead, ENEC significantly improves end-to-end inference performance, achieving up to a 6.3X speedup. On Ascend NPUs, ENEC is the first open-source lossless compression algorithm for model weights that achieves performance comparable to state-of-the-art GPU compressors, offering an effective solution for deploying large-scale AI models.
title ENEC: A Lossless AI Model Compression Method Enabling Fast Inference on Ascend NPUs
topic Hardware Architecture
Distributed, Parallel, and Cluster Computing
Machine Learning
url https://arxiv.org/abs/2604.03298