Post-Training Quantization of OpenPangu Models for Efficient Deployment on Atlas A2
Fuente:
arXiv
Guardado en:
| Autores principales: | Luo, Yilun, Zheng, Huaqing, Meng, Haoqian, Liu, Wenyuan, Zhang, Peng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MicroMix: Efficient Mixed-Precision Quantization with Microscaling Formats for Large Language Models
por: Liu, Wenyuan, et al.
Publicado: (2025)
por: Liu, Wenyuan, et al.
Publicado: (2025)
ARCQuant: Boosting NVFP4 Quantization with Augmented Residual Channels for LLMs
por: Meng, Haoqian, et al.
Publicado: (2026)
por: Meng, Haoqian, et al.
Publicado: (2026)
CrossQuant: A Post-Training Quantization Method with Smaller Quantization Kernel for Precise Large Language Model Compression
por: Liu, Wenyuan, et al.
Publicado: (2024)
por: Liu, Wenyuan, et al.
Publicado: (2024)
EfficientQAT: Efficient Quantization-Aware Training for Large Language Models
por: Chen, Mengzhao, et al.
Publicado: (2024)
por: Chen, Mengzhao, et al.
Publicado: (2024)
DAQ: Density-Aware Post-Training Weight-Only Quantization For LLMs
por: Luo, Yingsong, et al.
Publicado: (2024)
por: Luo, Yingsong, et al.
Publicado: (2024)
Beacon: Post-Training Quantization with Integrated Grid Selection
por: Zhang, Shihao, et al.
Publicado: (2025)
por: Zhang, Shihao, et al.
Publicado: (2025)
PTQTP: Post-Training Quantization to Trit-Planes for Large Language Models
por: Xiao, He, et al.
Publicado: (2025)
por: Xiao, He, et al.
Publicado: (2025)
Post Training Quantization of Large Language Models with Microscaling Formats
por: Sharify, Sayeh, et al.
Publicado: (2024)
por: Sharify, Sayeh, et al.
Publicado: (2024)
A Quantized VAE-MLP Botnet Detection Model: A Systematic Evaluation of Quantization-Aware Training and Post-Training Quantization Strategies
por: Wasswa, Hassan, et al.
Publicado: (2025)
por: Wasswa, Hassan, et al.
Publicado: (2025)
SLMQuant:Benchmarking Small Language Model Quantization for Practical Deployment
por: Wang, Jiacheng, et al.
Publicado: (2025)
por: Wang, Jiacheng, et al.
Publicado: (2025)
Q-realign: Piggybacking Realignment on Quantization for Safe and Efficient LLM Deployment
por: Tan, Qitao, et al.
Publicado: (2026)
por: Tan, Qitao, et al.
Publicado: (2026)
SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
por: Xiao, Guangxuan, et al.
Publicado: (2022)
por: Xiao, Guangxuan, et al.
Publicado: (2022)
D$^2$Quant: Accurate Low-bit Post-Training Weight Quantization for LLMs
por: Yan, Xianglong, et al.
Publicado: (2026)
por: Yan, Xianglong, et al.
Publicado: (2026)
Exploring Layer-wise Information Effectiveness for Post-Training Quantization in Small Language Models
por: Xiao, He, et al.
Publicado: (2025)
por: Xiao, He, et al.
Publicado: (2025)
Improving Quantization with Post-Training Model Expansion
por: Franco, Giuseppe, et al.
Publicado: (2025)
por: Franco, Giuseppe, et al.
Publicado: (2025)
RaanA: A Fast, Flexible, and Data-Efficient Post-Training Quantization Algorithm
por: Yang, Yongyi, et al.
Publicado: (2025)
por: Yang, Yongyi, et al.
Publicado: (2025)
AsymKV: Enabling 1-Bit Quantization of KV Cache with Layer-Wise Asymmetric Quantization Configurations
por: Tao, Qian, et al.
Publicado: (2024)
por: Tao, Qian, et al.
Publicado: (2024)
Efficient Edge LLMs Deployment via HessianAware Quantization and CPU GPU Collaborative
por: Zhang, Tuo, et al.
Publicado: (2025)
por: Zhang, Tuo, et al.
Publicado: (2025)
Accumulator-Aware Post-Training Quantization for Large Language Models
por: Colbert, Ian, et al.
Publicado: (2024)
por: Colbert, Ian, et al.
Publicado: (2024)
Pushing the Limits of Block Rotations in Post-Training Quantization
por: Sanjeet, Sai, et al.
Publicado: (2026)
por: Sanjeet, Sai, et al.
Publicado: (2026)
Benchmarking Post-Training Quantization in LLMs: Comprehensive Taxonomy, Unified Evaluation, and Comparative Analysis
por: Zhao, Jiaqi, et al.
Publicado: (2025)
por: Zhao, Jiaqi, et al.
Publicado: (2025)
MoBiE: Efficient Inference of Mixture of Binary Experts under Post-Training Quantization
por: Zhao, Zhixiong, et al.
Publicado: (2026)
por: Zhao, Zhixiong, et al.
Publicado: (2026)
Quamba: A Post-Training Quantization Recipe for Selective State Space Models
por: Chiang, Hung-Yueh, et al.
Publicado: (2024)
por: Chiang, Hung-Yueh, et al.
Publicado: (2024)
Achieving binary weight and activation for LLMs using Post-Training Quantization
por: Song, Siqing, et al.
Publicado: (2025)
por: Song, Siqing, et al.
Publicado: (2025)
DAQ: Delta-Aware Quantization for Post-Training LLM Weight Compression
por: Yu, Xiaoming, et al.
Publicado: (2026)
por: Yu, Xiaoming, et al.
Publicado: (2026)
SmallThinker: A Family of Efficient Large Language Models Natively Trained for Local Deployment
por: Song, Yixin, et al.
Publicado: (2025)
por: Song, Yixin, et al.
Publicado: (2025)
MagR: Weight Magnitude Reduction for Enhancing Post-Training Quantization
por: Zhang, Aozhong, et al.
Publicado: (2024)
por: Zhang, Aozhong, et al.
Publicado: (2024)
Quant-dLLM: Post-Training Extreme Low-Bit Quantization for Diffusion Large Language Models
por: Zhang, Tianao, et al.
Publicado: (2025)
por: Zhang, Tianao, et al.
Publicado: (2025)
BiLLM: Pushing the Limit of Post-Training Quantization for LLMs
por: Huang, Wei, et al.
Publicado: (2024)
por: Huang, Wei, et al.
Publicado: (2024)
Activation Sensitivity as a Unifying Principle for Post-Training Quantization
por: Xu, Bruce Changlong
Publicado: (2026)
por: Xu, Bruce Changlong
Publicado: (2026)
Qronos: Correcting the Past by Shaping the Future... in Post-Training Quantization
por: Zhang, Shihao, et al.
Publicado: (2025)
por: Zhang, Shihao, et al.
Publicado: (2025)
Towards Next-Level Post-Training Quantization of Hyper-Scale Transformers
por: Kim, Junhan, et al.
Publicado: (2024)
por: Kim, Junhan, et al.
Publicado: (2024)
Post-training for Efficient Communication via Convention Formation
por: Hua, Yilun, et al.
Publicado: (2025)
por: Hua, Yilun, et al.
Publicado: (2025)
Advancing Model Refinement: Muon-Optimized Distillation and Quantization for LLM Deployment
por: Sander, Jacob, et al.
Publicado: (2026)
por: Sander, Jacob, et al.
Publicado: (2026)
PTQ1.61: Push the Real Limit of Extremely Low-Bit Post-Training Quantization Methods for Large Language Models
por: Zhao, Jiaqi, et al.
Publicado: (2025)
por: Zhao, Jiaqi, et al.
Publicado: (2025)
Resource-Efficient Generative AI Model Deployment in Mobile Edge Networks
por: Liang, Yuxin, et al.
Publicado: (2024)
por: Liang, Yuxin, et al.
Publicado: (2024)
DiRL: An Efficient Post-Training Framework for Diffusion Language Models
por: Zhu, Ying, et al.
Publicado: (2025)
por: Zhu, Ying, et al.
Publicado: (2025)
Learning-Zone Energy: Online Data Selection for Efficient RL Post-Training
por: Cui, Peng, et al.
Publicado: (2026)
por: Cui, Peng, et al.
Publicado: (2026)
Interactions Across Blocks in Post-Training Quantization of Large Language Models
por: Shabanovi, Khasmamad, et al.
Publicado: (2024)
por: Shabanovi, Khasmamad, et al.
Publicado: (2024)
Q-Palette: Fractional-Bit Quantizers Toward Optimal Bit Allocation for Efficient LLM Deployment
por: Lee, Deokjae, et al.
Publicado: (2025)
por: Lee, Deokjae, et al.
Publicado: (2025)
Ejemplares similares
-
MicroMix: Efficient Mixed-Precision Quantization with Microscaling Formats for Large Language Models
por: Liu, Wenyuan, et al.
Publicado: (2025) -
ARCQuant: Boosting NVFP4 Quantization with Augmented Residual Channels for LLMs
por: Meng, Haoqian, et al.
Publicado: (2026) -
CrossQuant: A Post-Training Quantization Method with Smaller Quantization Kernel for Precise Large Language Model Compression
por: Liu, Wenyuan, et al.
Publicado: (2024) -
EfficientQAT: Efficient Quantization-Aware Training for Large Language Models
por: Chen, Mengzhao, et al.
Publicado: (2024) -
DAQ: Density-Aware Post-Training Weight-Only Quantization For LLMs
por: Luo, Yingsong, et al.
Publicado: (2024)