ENEC: A Lossless AI Model Compression Method Enabling Fast Inference on Ascend NPUs
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Jinwu, Wu, Jiaan, Liu, Zedong, Ma, Xinyang, Zhao, Hairui, Gu, Yida, Huang, Yuanhong, Liu, Xingchen, Huang, Wenjing, Wei, Zheng, Xing, Jing, Ma, Yili, Zhang, Qingyi, An, Baoyi, Hu, Zhongzhe, Liu, Shaoteng, Zhu, Xia, Lu, Jiaxun, Tan, Guangming, Tao, Dingwen |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Benchmarking Ultra-Low-Power $μ$NPUs
by: Millar, Josh, et al.
Published: (2025)
by: Millar, Josh, et al.
Published: (2025)
TRACE: Unlocking Effective CXL Bandwidth via Lossless Compression and Precision Scaling
by: Xie, Rui, et al.
Published: (2025)
by: Xie, Rui, et al.
Published: (2025)
KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving
by: Liu, Zedong, et al.
Published: (2026)
by: Liu, Zedong, et al.
Published: (2026)
NVR: Vector Runahead on NPUs for Sparse Memory Access
by: Wang, Hui, et al.
Published: (2025)
by: Wang, Hui, et al.
Published: (2025)
T-MAN: Enabling End-to-End Low-Bit LLM Inference on NPUs via Unified Table Lookup
by: Wei, Jianyu, et al.
Published: (2025)
by: Wei, Jianyu, et al.
Published: (2025)
MemExplorer: Navigating the Heterogeneous Memory Design Space for Agentic Inference NPUs
by: Wu, Haoran, et al.
Published: (2026)
by: Wu, Haoran, et al.
Published: (2026)
Striking the Balance: GEMM Performance Optimization Across Generations of Ryzen AI NPUs
by: Taka, Endri, et al.
Published: (2025)
by: Taka, Endri, et al.
Published: (2025)
Layer-wise Weight Selection for Power-Efficient Neural Network Acceleration
by: Fang, Jiaxun, et al.
Published: (2025)
by: Fang, Jiaxun, et al.
Published: (2025)
VeriCache: Turning Lossy KV Cache into Lossless LLM Inference
by: Yao, Jiayi, et al.
Published: (2026)
by: Yao, Jiayi, et al.
Published: (2026)
Ascend HiFloat8 Format for Deep Learning
by: Luo, Yuanyong, et al.
Published: (2024)
by: Luo, Yuanyong, et al.
Published: (2024)
From Principles to Practice: A Systematic Study of LLM Serving on Multi-core NPUs
by: Zhu, Tianhao, et al.
Published: (2025)
by: Zhu, Tianhao, et al.
Published: (2025)
When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference
by: Li, Pu, et al.
Published: (2026)
by: Li, Pu, et al.
Published: (2026)
SHIELD: A Segmented Hierarchical Memory Architecture for Energy-Efficient LLM Inference on Edge NPUs
by: Zhang, Jintao, et al.
Published: (2026)
by: Zhang, Jintao, et al.
Published: (2026)
CaMDN: Enhancing Cache Efficiency for Multi-tenant DNNs on Integrated NPUs
by: Cai, Tianhao, et al.
Published: (2025)
by: Cai, Tianhao, et al.
Published: (2025)
LEXI: Lossless Exponent Coding for Efficient Inter-Chiplet Communication in Hybrid LLMs
by: Sun, Miao, et al.
Published: (2026)
by: Sun, Miao, et al.
Published: (2026)
RAS: A Bit-Exact rANS Accelerator For High-Performance Neural Lossless Compression
by: Qin, Yuchao, et al.
Published: (2025)
by: Qin, Yuchao, et al.
Published: (2025)
From Loop Nests to Silicon: Mapping AI Workloads onto AMD NPUs with MLIR-AIR
by: Wang, Erwei, et al.
Published: (2025)
by: Wang, Erwei, et al.
Published: (2025)
Splatonic: Architecture Support for 3D Gaussian Splatting SLAM via Sparse Processing
by: Huang, Xiaotong, et al.
Published: (2025)
by: Huang, Xiaotong, et al.
Published: (2025)
CIM-Tuner: Balancing the Compute and Storage Capacity of SRAM-CIM Accelerator via Hardware-mapping Co-exploration
by: Chen, Jinwu, et al.
Published: (2026)
by: Chen, Jinwu, et al.
Published: (2026)
Scope: A Scalable Merged Pipeline Framework for Multi-Chip-Module NN Accelerators
by: Huang, Zongle, et al.
Published: (2026)
by: Huang, Zongle, et al.
Published: (2026)
HFRWKV: A High-Performance Fully On-Chip Hardware Accelerator for RWKV
by: Shijie, Liu, et al.
Published: (2026)
by: Shijie, Liu, et al.
Published: (2026)
Algorithm-hardware co-design for Energy-Efficient A/D conversion in ReRAM-based accelerators
by: Zhang, Chenguang, et al.
Published: (2024)
by: Zhang, Chenguang, et al.
Published: (2024)
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference
by: Yubeaton, Patrick, et al.
Published: (2025)
by: Yubeaton, Patrick, et al.
Published: (2025)
RLPlanner: Reinforcement Learning based Floorplanning for Chiplets with Fast Thermal Analysis
by: Duan, Yuanyuan, et al.
Published: (2023)
by: Duan, Yuanyuan, et al.
Published: (2023)
Reducing the Cost of Dropout in Flash-Attention by Hiding RNG with GEMM
by: Ma, Haiyue, et al.
Published: (2024)
by: Ma, Haiyue, et al.
Published: (2024)
Reimagining Memory Access for LLM Inference: Compression-Aware Memory Controller Design
by: Xie, Rui, et al.
Published: (2025)
by: Xie, Rui, et al.
Published: (2025)
DaDu-Corki: Algorithm-Architecture Co-Design for Embodied AI-powered Robotic Manipulation
by: Huang, Yiyang, et al.
Published: (2024)
by: Huang, Yiyang, et al.
Published: (2024)
Breaking the HBM Bit Cost Barrier: Domain-Specific ECC for AI Inference Infrastructure
by: Xie, Rui, et al.
Published: (2025)
by: Xie, Rui, et al.
Published: (2025)
Making Strong Error-Correcting Codes Work Effectively for HBM in AI Inference
by: Xie, Rui, et al.
Published: (2025)
by: Xie, Rui, et al.
Published: (2025)
31.1 A 14.08-to-135.69Token/s ReRAM-on-Logic Stacked Outlier-Free Large-Language-Model Accelerator with Block-Clustered Weight-Compression and Adaptive Parallel-Speculative-Decoding
by: Dong, Pingcheng, et al.
Published: (2026)
by: Dong, Pingcheng, et al.
Published: (2026)
Wit-HW: Bug Localization in Hardware Design Code via Witness Test Case Generation
by: Ma, Ruiyang, et al.
Published: (2025)
by: Ma, Ruiyang, et al.
Published: (2025)
A Novel Computing Paradigm for MobileNetV3 using Memristor
by: Li, Jiale, et al.
Published: (2024)
by: Li, Jiale, et al.
Published: (2024)
VeRA+: Vector-Based Lightweight Digital Compensation for Drift-Resilient RRAM In-Memory Computing
by: Dong, Weirong, et al.
Published: (2026)
by: Dong, Weirong, et al.
Published: (2026)
ML-based AIG Timing Prediction to Enhance Logic Optimization
by: Jiang, Wenjing, et al.
Published: (2024)
by: Jiang, Wenjing, et al.
Published: (2024)
An O(m+n)-Space Spatiotemporal Denoising Filter with Cache-Like Memories for Dynamic Vision Sensors
by: Zhao, Qinghang, et al.
Published: (2024)
by: Zhao, Qinghang, et al.
Published: (2024)
Hecaton: Training Large Language Models with Scalable Chiplet Systems
by: Huang, Zongle, et al.
Published: (2024)
by: Huang, Zongle, et al.
Published: (2024)
DRACO: Co-design for DSP-Efficient Rigid Body Dynamics Accelerator
by: Liu, Xingyu, et al.
Published: (2025)
by: Liu, Xingyu, et al.
Published: (2025)
RPCAcc: A High-Performance and Reconfigurable PCIe-attached RPC Accelerator
by: Zhang, Jie, et al.
Published: (2024)
by: Zhang, Jie, et al.
Published: (2024)
A Sparsity-Aware Autonomous Path Planning Accelerator with HW/SW Co-Design and Multi-Level Dataflow Optimization
by: Zhang, Yifan, et al.
Published: (2025)
by: Zhang, Yifan, et al.
Published: (2025)
TENET: An Efficient Sparsity-Aware LUT-Centric Architecture for Ternary LLM Inference On Edge
by: Huang, Zhirui, et al.
Published: (2025)
by: Huang, Zhirui, et al.
Published: (2025)
Similar Items
-
Benchmarking Ultra-Low-Power $μ$NPUs
by: Millar, Josh, et al.
Published: (2025) -
TRACE: Unlocking Effective CXL Bandwidth via Lossless Compression and Precision Scaling
by: Xie, Rui, et al.
Published: (2025) -
KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving
by: Liu, Zedong, et al.
Published: (2026) -
NVR: Vector Runahead on NPUs for Sparse Memory Access
by: Wang, Hui, et al.
Published: (2025) -
T-MAN: Enabling End-to-End Low-Bit LLM Inference on NPUs via Unified Table Lookup
by: Wei, Jianyu, et al.
Published: (2025)