Beyond Outliers: A Data-Free Layer-wise Mixed-Precision Quantization Approach Driven by Numerical and Structural Dual-Sensitivity
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Hengyuan, Chen, Xinrong, Su, Zunhai, Liang, Xiao, Xiong, Jing, Xu, Wendong, Xiao, He, Tao, Chaofan, Zhang, Wei, Xie, Ruobing, Jiang, Lei, So, Hayden Kwok-Hay, Wong, Ngai |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Exploring Layer-wise Information Effectiveness for Post-Training Quantization in Small Language Models
di: Xiao, He, et al.
Pubblicazione: (2025)
di: Xiao, He, et al.
Pubblicazione: (2025)
Find Your Optimal Teacher: Personalized Data Synthesis via Router-Guided Multi-Teacher Distillation
di: Zhang, Hengyuan, et al.
Pubblicazione: (2025)
di: Zhang, Hengyuan, et al.
Pubblicazione: (2025)
GuiLoMo: Allocating Expert Number and Rank for LoRA-MoE via Bilevel Optimization with GuidedSelection Vectors
di: Zhang, Hengyuan, et al.
Pubblicazione: (2025)
di: Zhang, Hengyuan, et al.
Pubblicazione: (2025)
DyBit: Dynamic Bit-Precision Numbers for Efficient Quantized Neural Network Inference
di: Zhou, Jiajun, et al.
Pubblicazione: (2023)
di: Zhou, Jiajun, et al.
Pubblicazione: (2023)
Residual Decoding: Mitigating Hallucinations in Large Vision-Language Models via History-Aware Residual Guidance
di: Chen, Xinrong, et al.
Pubblicazione: (2026)
di: Chen, Xinrong, et al.
Pubblicazione: (2026)
CktFormalizer: Autoformalization of Natural Language into Circuit Representations
di: Xiong, Jing, et al.
Pubblicazione: (2026)
di: Xiong, Jing, et al.
Pubblicazione: (2026)
TreeReview: A Dynamic Tree of Questions Framework for Deep and Efficient LLM-based Scientific Peer Review
di: Chang, Yuan, et al.
Pubblicazione: (2025)
di: Chang, Yuan, et al.
Pubblicazione: (2025)
TATAA: Programmable Mixed-Precision Transformer Acceleration with a Transformable Arithmetic Architecture
di: Wu, Jiajun, et al.
Pubblicazione: (2024)
di: Wu, Jiajun, et al.
Pubblicazione: (2024)
A Composable Dynamic Sparse Dataflow Architecture for Efficient Event-based Vision Processing on FPGA
di: Gao, Yizhao, et al.
Pubblicazione: (2024)
di: Gao, Yizhao, et al.
Pubblicazione: (2024)
Co-designing a Sub-millisecond Latency Event-based Eye Tracking System with Submanifold Sparse CNN
di: Zhang, Baoheng, et al.
Pubblicazione: (2024)
di: Zhang, Baoheng, et al.
Pubblicazione: (2024)
Mix-QViT: Mixed-Precision Vision Transformer Quantization Driven by Layer Importance and Quantization Sensitivity
di: Ranjan, Navin, et al.
Pubblicazione: (2025)
di: Ranjan, Navin, et al.
Pubblicazione: (2025)
Layer-wise Quantization for Quantized Optimistic Dual Averaging
di: Nguyen, Anh Duc, et al.
Pubblicazione: (2025)
di: Nguyen, Anh Duc, et al.
Pubblicazione: (2025)
LampQ: Towards Accurate Layer-wise Mixed Precision Quantization for Vision Transformers
di: Kim, Minjun, et al.
Pubblicazione: (2025)
di: Kim, Minjun, et al.
Pubblicazione: (2025)
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models
di: Zhang, Hengyuan, et al.
Pubblicazione: (2026)
di: Zhang, Hengyuan, et al.
Pubblicazione: (2026)
APTQ: Attention-aware Post-Training Mixed-Precision Quantization for Large Language Models
di: Guan, Ziyi, et al.
Pubblicazione: (2024)
di: Guan, Ziyi, et al.
Pubblicazione: (2024)
LRP-QViT: Mixed-Precision Vision Transformer Quantization via Layer-wise Relevance Propagation
di: Ranjan, Navin, et al.
Pubblicazione: (2024)
di: Ranjan, Navin, et al.
Pubblicazione: (2024)
AnchorTP: Resilient LLM Inference with State-Preserving Elastic Tensor Parallelism
di: Xu, Wendong, et al.
Pubblicazione: (2025)
di: Xu, Wendong, et al.
Pubblicazione: (2025)
KVSink: Understanding and Enhancing the Preservation of Attention Sinks in KV Cache Quantization for LLMs
di: Su, Zunhai, et al.
Pubblicazione: (2025)
di: Su, Zunhai, et al.
Pubblicazione: (2025)
SpikeMOT: Event-based Multi-Object Tracking with Sparse Motion Features
di: Wang, Song, et al.
Pubblicazione: (2023)
di: Wang, Song, et al.
Pubblicazione: (2023)
PTQTP: Post-Training Quantization to Trit-Planes for Large Language Models
di: Xiao, He, et al.
Pubblicazione: (2025)
di: Xiao, He, et al.
Pubblicazione: (2025)
DoPE: Denoising Rotary Position Embedding
di: Xiong, Jing, et al.
Pubblicazione: (2025)
di: Xiong, Jing, et al.
Pubblicazione: (2025)
RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations
di: Su, Zunhai, et al.
Pubblicazione: (2025)
di: Su, Zunhai, et al.
Pubblicazione: (2025)
BATQuant: Outlier-resilient MXFP4 Quantization via Learnable Block-wise Optimization
di: Li, Ji-Fu, et al.
Pubblicazione: (2026)
di: Li, Ji-Fu, et al.
Pubblicazione: (2026)
MUXQ: Mixed-to-Uniform Precision MatriX Quantization via Low-Rank Outlier Decomposition
di: Lee, Seoungsub, et al.
Pubblicazione: (2026)
di: Lee, Seoungsub, et al.
Pubblicazione: (2026)
Fully Integrated Memristive Spiking Neural Network with Analog Neurons for High-Speed Event-Based Data Processing
di: Wang, Zhu, et al.
Pubblicazione: (2025)
di: Wang, Zhu, et al.
Pubblicazione: (2025)
Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation
di: Su, Zunhai, et al.
Pubblicazione: (2026)
di: Su, Zunhai, et al.
Pubblicazione: (2026)
Joint Pruning and Channel-wise Mixed-Precision Quantization for Efficient Deep Neural Networks
di: Motetti, Beatrice Alessandra, et al.
Pubblicazione: (2024)
di: Motetti, Beatrice Alessandra, et al.
Pubblicazione: (2024)
Mixed-Precision Graph Neural Quantization for Low Bit Large Language Models
di: Liu, Wanlong, et al.
Pubblicazione: (2025)
di: Liu, Wanlong, et al.
Pubblicazione: (2025)
KVTuner: Sensitivity-Aware Layer-Wise Mixed-Precision KV Cache Quantization for Efficient and Nearly Lossless LLM Inference
di: Li, Xing, et al.
Pubblicazione: (2025)
di: Li, Xing, et al.
Pubblicazione: (2025)
Beyond Outliers: A Study of Optimizers Under Quantization
di: Vlassis, Georgios, et al.
Pubblicazione: (2025)
di: Vlassis, Georgios, et al.
Pubblicazione: (2025)
XStreamVGGT: Extremely Memory-Efficient Streaming Vision Geometry Grounded Transformer with KV Cache Compression
di: Su, Zunhai, et al.
Pubblicazione: (2026)
di: Su, Zunhai, et al.
Pubblicazione: (2026)
XStreamVGGT: Extremely Memory-Efficient Streaming Vision Geometry Grounded Transformer with KV Cache Compression
di: Su, Zunhai, et al.
Pubblicazione: (2026)
di: Su, Zunhai, et al.
Pubblicazione: (2026)
OMPQ: Orthogonal Mixed Precision Quantization
di: Ma, Yuexiao, et al.
Pubblicazione: (2021)
di: Ma, Yuexiao, et al.
Pubblicazione: (2021)
Sensitivity-Aware Mixed-Precision Quantization for ReRAM-based Computing-in-Memory
di: Chen, Guan-Cheng, et al.
Pubblicazione: (2025)
di: Chen, Guan-Cheng, et al.
Pubblicazione: (2025)
Lifecycle Cost-Effectiveness Modeling for Redundancy-Enhanced Multi-Chiplet Architectures
di: Liu, Zizhen, et al.
Pubblicazione: (2026)
di: Liu, Zizhen, et al.
Pubblicazione: (2026)
BPDQ: Bit-Plane Decomposition Quantization on a Variable Grid for Large Language Models
di: Chen, Junyu, et al.
Pubblicazione: (2026)
di: Chen, Junyu, et al.
Pubblicazione: (2026)
LoaQ: Layer-wise Output Approximation Quantization
di: Lin, Li, et al.
Pubblicazione: (2025)
di: Lin, Li, et al.
Pubblicazione: (2025)
Evaluating Numerical Accuracy in Mixed-Precision Computing by Dual-Delta Testing
di: Xie, Peichen
Pubblicazione: (2026)
di: Xie, Peichen
Pubblicazione: (2026)
CodeQuant: Unified Clustering and Quantization for Enhanced Outlier Smoothing in Low-Precision Mixture-of-Experts
di: Yin, Xiangyang, et al.
Pubblicazione: (2026)
di: Yin, Xiangyang, et al.
Pubblicazione: (2026)
Value-Driven Mixed-Precision Quantization for Patch-Based Inference on Microcontrollers
di: Tao, Wei, et al.
Pubblicazione: (2024)
di: Tao, Wei, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Exploring Layer-wise Information Effectiveness for Post-Training Quantization in Small Language Models
di: Xiao, He, et al.
Pubblicazione: (2025) -
Find Your Optimal Teacher: Personalized Data Synthesis via Router-Guided Multi-Teacher Distillation
di: Zhang, Hengyuan, et al.
Pubblicazione: (2025) -
GuiLoMo: Allocating Expert Number and Rank for LoRA-MoE via Bilevel Optimization with GuidedSelection Vectors
di: Zhang, Hengyuan, et al.
Pubblicazione: (2025) -
DyBit: Dynamic Bit-Precision Numbers for Efficient Quantized Neural Network Inference
di: Zhou, Jiajun, et al.
Pubblicazione: (2023) -
Residual Decoding: Mitigating Hallucinations in Large Vision-Language Models via History-Aware Residual Guidance
di: Chen, Xinrong, et al.
Pubblicazione: (2026)