Unleashing Low-Bit Inference on Ascend NPUs: A Comprehensive Evaluation of HiFloat Formats
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, Pengxiang, Zhen, Hui-Ling, Li, Xing, Bao, Han, Lin, Weizhe, Yang, Zhiyuan, Zhang, Manyi, Luo, Yuanyong, Yu, Ziwei, Wang, Xin, Yuan, Mingxuan, Yu, Xianzhi, Dong, Zhenhua |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HiFloat4 Format for Language Model Pre-training on Ascend NPUs
von: Taghian, Mehran, et al.
Veröffentlicht: (2026)
von: Taghian, Mehran, et al.
Veröffentlicht: (2026)
Ascend HiFloat8 Format for Deep Learning
von: Luo, Yuanyong, et al.
Veröffentlicht: (2024)
von: Luo, Yuanyong, et al.
Veröffentlicht: (2024)
HiFloat4 Format for Language Model Inference
von: Luo, Yuanyong, et al.
Veröffentlicht: (2026)
von: Luo, Yuanyong, et al.
Veröffentlicht: (2026)
Benchmarking Post-Training Quantization of Large Language Models under Microscaling Floating Point Formats
von: Zhang, Manyi, et al.
Veröffentlicht: (2026)
von: Zhang, Manyi, et al.
Veröffentlicht: (2026)
What Matters For Safety Alignment?
von: Li, Xing, et al.
Veröffentlicht: (2026)
von: Li, Xing, et al.
Veröffentlicht: (2026)
Tail-Aware HiFloat4: W4A4 Post-Training Quantization for Wan2.2
von: Feng, Zhanfeng, et al.
Veröffentlicht: (2026)
von: Feng, Zhanfeng, et al.
Veröffentlicht: (2026)
ENEC: A Lossless AI Model Compression Method Enabling Fast Inference on Ascend NPUs
von: Yang, Jinwu, et al.
Veröffentlicht: (2026)
von: Yang, Jinwu, et al.
Veröffentlicht: (2026)
Boundary-Protection W8A8 HiFloat8 Quantization for Large-Scale Text-to-Video Diffusion Transformers
von: Zhao, Yiming
Veröffentlicht: (2026)
von: Zhao, Yiming
Veröffentlicht: (2026)
SGEMM-cube: Precision-Recovery FP32 GEMM Approximation on Ascend NPUs with FP16 Matrix Engines
von: Xue, Weicheng, et al.
Veröffentlicht: (2025)
von: Xue, Weicheng, et al.
Veröffentlicht: (2025)
PreMoE: Proactive Inference for Efficient Mixture-of-Experts
von: Pei, Zehua, et al.
Veröffentlicht: (2025)
von: Pei, Zehua, et al.
Veröffentlicht: (2025)
TrimR: Verifier-based Training-Free Thinking Compression for Efficient Test-Time Scaling
von: Lin, Weizhe, et al.
Veröffentlicht: (2025)
von: Lin, Weizhe, et al.
Veröffentlicht: (2025)
Behavioral Fingerprinting of Large Language Models
von: Pei, Zehua, et al.
Veröffentlicht: (2025)
von: Pei, Zehua, et al.
Veröffentlicht: (2025)
What Makes Low-Bit Quantization-Aware Training Work for Reasoning LLMs? A Systematic Study
von: Lv, Keyu, et al.
Veröffentlicht: (2026)
von: Lv, Keyu, et al.
Veröffentlicht: (2026)
T-MAN: Enabling End-to-End Low-Bit LLM Inference on NPUs via Unified Table Lookup
von: Wei, Jianyu, et al.
Veröffentlicht: (2025)
von: Wei, Jianyu, et al.
Veröffentlicht: (2025)
Pangu Ultra MoE: How to Train Your Big MoE on Ascend NPUs
von: Tang, Yehui, et al.
Veröffentlicht: (2025)
von: Tang, Yehui, et al.
Veröffentlicht: (2025)
Pangu Ultra: Pushing the Limits of Dense Large Language Models on Ascend NPUs
von: Yin, Yichun, et al.
Veröffentlicht: (2025)
von: Yin, Yichun, et al.
Veröffentlicht: (2025)
SwiftMem: Fast Agentic Memory via Query-aware Indexing
von: Tian, Anxin, et al.
Veröffentlicht: (2026)
von: Tian, Anxin, et al.
Veröffentlicht: (2026)
EAGLE-Pangu: Accelerator-Safe Tree Speculative Decoding on Ascend NPUs
von: Han, Chang, et al.
Veröffentlicht: (2026)
von: Han, Chang, et al.
Veröffentlicht: (2026)
Towards Efficient Agents: A Co-Design of Inference Architecture and System
von: Lin, Weizhe, et al.
Veröffentlicht: (2025)
von: Lin, Weizhe, et al.
Veröffentlicht: (2025)
BATQuant: Outlier-resilient MXFP4 Quantization via Learnable Block-wise Optimization
von: Li, Ji-Fu, et al.
Veröffentlicht: (2026)
von: Li, Ji-Fu, et al.
Veröffentlicht: (2026)
MindVL: Towards Efficient and Effective Training of Multimodal Large Language Models on Ascend NPUs
von: Chen, Feilong, et al.
Veröffentlicht: (2025)
von: Chen, Feilong, et al.
Veröffentlicht: (2025)
Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling
von: Sun, Shengyin, et al.
Veröffentlicht: (2025)
von: Sun, Shengyin, et al.
Veröffentlicht: (2025)
Enhancing Learned Knowledge in LoRA Adapters Through Efficient Contrastive Decoding on Ascend NPUs
von: Heisler, Morgan Lindsay, et al.
Veröffentlicht: (2025)
von: Heisler, Morgan Lindsay, et al.
Veröffentlicht: (2025)
SVDq: 1.25-bit and 410x Key Cache Compression for LLM Attention
von: Yankun, Hong, et al.
Veröffentlicht: (2025)
von: Yankun, Hong, et al.
Veröffentlicht: (2025)
Revisiting Judge Decoding from First Principles via Training-Free Distributional Divergence
von: Sun, Shengyin, et al.
Veröffentlicht: (2026)
von: Sun, Shengyin, et al.
Veröffentlicht: (2026)
FastAttention: Extend FlashAttention2 to NPUs and Low-resource GPUs
von: Lin, Haoran, et al.
Veröffentlicht: (2024)
von: Lin, Haoran, et al.
Veröffentlicht: (2024)
Quant-Trim in Practice: Improved Cross-Platform Low-Bit Deployment on Edge NPUs
von: Dhahri, Rayen, et al.
Veröffentlicht: (2025)
von: Dhahri, Rayen, et al.
Veröffentlicht: (2025)
Multi‐Bit Floating‐Gate Memory with an Ultrawide Programmable Window
von: Ce Li, et al.
Veröffentlicht: (2026)
von: Ce Li, et al.
Veröffentlicht: (2026)
FocuSFT: Bilevel Optimization for Dilution-Aware Long-Context Fine-Tuning
von: Pei, Zehua, et al.
Veröffentlicht: (2026)
von: Pei, Zehua, et al.
Veröffentlicht: (2026)
From Pruning to Grafting: Dynamic Knowledge Redistribution via Learnable Layer Fusion
von: Pei, Zehua, et al.
Veröffentlicht: (2024)
von: Pei, Zehua, et al.
Veröffentlicht: (2024)
Fast On-device LLM Inference with NPUs
von: Xu, Daliang, et al.
Veröffentlicht: (2024)
von: Xu, Daliang, et al.
Veröffentlicht: (2024)
PSD: Pushing the Pareto Frontier of Diffusion LLMs via Parallel Speculative Decoding
von: Sun, Shengyin, et al.
Veröffentlicht: (2026)
von: Sun, Shengyin, et al.
Veröffentlicht: (2026)
QuantClaw: Precision Where It Matters for OpenClaw
von: Zhang, Manyi, et al.
Veröffentlicht: (2026)
von: Zhang, Manyi, et al.
Veröffentlicht: (2026)
NITRO: LLM Inference on Intel Laptop NPUs
von: Fei, Anthony, et al.
Veröffentlicht: (2024)
von: Fei, Anthony, et al.
Veröffentlicht: (2024)
OSUM-Pangu: An Open-Source Multidimension Speech Understanding Foundation Model Built upon OpenPangu on Ascend NPUs
von: Liao, Yujie, et al.
Veröffentlicht: (2026)
von: Liao, Yujie, et al.
Veröffentlicht: (2026)
Benchmarking Ultra-Low-Power $μ$NPUs
von: Millar, Josh, et al.
Veröffentlicht: (2025)
von: Millar, Josh, et al.
Veröffentlicht: (2025)
EPD-Serve: A Flexible Multimodal EPD Disaggregation Inference Serving System On Ascend
von: Bai, Fan, et al.
Veröffentlicht: (2026)
von: Bai, Fan, et al.
Veröffentlicht: (2026)
W4A16 Mixed-Precision Matrix Multiplication on Decoupled Architecture: Kernel Design and Memory Bottleneck Analysis for Ascend NPUs
von: He, Yuanhong, et al.
Veröffentlicht: (2026)
von: He, Yuanhong, et al.
Veröffentlicht: (2026)
Analytical FFN-to-MoE Restructuring via Activation Pattern Analysis
von: Pei, Zehua, et al.
Veröffentlicht: (2025)
von: Pei, Zehua, et al.
Veröffentlicht: (2025)
Efficient Mixture-of-Experts LLM Inference with Apple Silicon NPUs
von: Benazir, Afsara, et al.
Veröffentlicht: (2026)
von: Benazir, Afsara, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
HiFloat4 Format for Language Model Pre-training on Ascend NPUs
von: Taghian, Mehran, et al.
Veröffentlicht: (2026) -
Ascend HiFloat8 Format for Deep Learning
von: Luo, Yuanyong, et al.
Veröffentlicht: (2024) -
HiFloat4 Format for Language Model Inference
von: Luo, Yuanyong, et al.
Veröffentlicht: (2026) -
Benchmarking Post-Training Quantization of Large Language Models under Microscaling Floating Point Formats
von: Zhang, Manyi, et al.
Veröffentlicht: (2026) -
What Matters For Safety Alignment?
von: Li, Xing, et al.
Veröffentlicht: (2026)