EdgeRazor: A Lightweight Framework for Large Language Models via Mixed-Precision Quantization-Aware Distillation
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Shu-Hao, Huang, Le-Tong, Deng, Xiang-Sheng, Zou, Xin-Yi, Wu, Chen, Li, Nan, Zhang, Shao-Qun, Zhou, Zhi-Hua |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
LQA: A Lightweight Quantized-Adaptive Framework for Vision-Language Models on the Edge
di: Wang, Xin, et al.
Pubblicazione: (2026)
di: Wang, Xin, et al.
Pubblicazione: (2026)
APreQEL: Adaptive Mixed Precision Quantization For Edge LLMs
di: Bouzouad, Meriem, et al.
Pubblicazione: (2026)
di: Bouzouad, Meriem, et al.
Pubblicazione: (2026)
MixKVQ: Query-Aware Mixed-Precision KV Cache Quantization for Long-Context Reasoning
di: Zhang, Tao, et al.
Pubblicazione: (2025)
di: Zhang, Tao, et al.
Pubblicazione: (2025)
AMED: Automatic Mixed-Precision Quantization for Edge Devices
di: Kimhi, Moshe, et al.
Pubblicazione: (2022)
di: Kimhi, Moshe, et al.
Pubblicazione: (2022)
Generalization Bounds of Spiking Neural Networks via Rademacher Complexity
di: Zhang, Shao-Qun, et al.
Pubblicazione: (2026)
di: Zhang, Shao-Qun, et al.
Pubblicazione: (2026)
On-Chip Hardware-Aware Quantization for Mixed Precision Neural Networks
di: Huang, Wei, et al.
Pubblicazione: (2023)
di: Huang, Wei, et al.
Pubblicazione: (2023)
Mixed-Precision Quantization for Deep Vision Models with Integer Quadratic Programming
di: Deng, Zihao, et al.
Pubblicazione: (2023)
di: Deng, Zihao, et al.
Pubblicazione: (2023)
Sensitivity-Aware Mixed-Precision Quantization for ReRAM-based Computing-in-Memory
di: Chen, Guan-Cheng, et al.
Pubblicazione: (2025)
di: Chen, Guan-Cheng, et al.
Pubblicazione: (2025)
TernaryCLIP: Efficiently Compressing Vision-Language Models with Ternary Weights and Distilled Knowledge
di: Zhang, Shu-Hao, et al.
Pubblicazione: (2025)
di: Zhang, Shu-Hao, et al.
Pubblicazione: (2025)
OMPQ: Orthogonal Mixed Precision Quantization
di: Ma, Yuexiao, et al.
Pubblicazione: (2021)
di: Ma, Yuexiao, et al.
Pubblicazione: (2021)
TinyMusician: On-Device Music Generation with Knowledge Distillation and Mixed Precision Quantization
di: Wang, Hainan, et al.
Pubblicazione: (2025)
di: Wang, Hainan, et al.
Pubblicazione: (2025)
MoQAE: Mixed-Precision Quantization for Long-Context LLM Inference via Mixture of Quantization-Aware Experts
di: Tao, Wei, et al.
Pubblicazione: (2025)
di: Tao, Wei, et al.
Pubblicazione: (2025)
FLIQS: One-Shot Mixed-Precision Floating-Point and Integer Quantization Search
di: Dotzel, Jordan, et al.
Pubblicazione: (2023)
di: Dotzel, Jordan, et al.
Pubblicazione: (2023)
Learning on a Razor's Edge: Identifiability and Singularity of Polynomial Neural Networks
di: Shahverdi, Vahid, et al.
Pubblicazione: (2025)
di: Shahverdi, Vahid, et al.
Pubblicazione: (2025)
Qrazor: Reliable and Effortless 4-bit LLM Quantization by Significant Data Razoring
di: Lee, Dongyoung, et al.
Pubblicazione: (2025)
di: Lee, Dongyoung, et al.
Pubblicazione: (2025)
Adaptive Distribution-aware Quantization for Mixed-Precision Neural Networks
di: Jia, Shaohang, et al.
Pubblicazione: (2025)
di: Jia, Shaohang, et al.
Pubblicazione: (2025)
FairQuant: Fairness-Aware Mixed-Precision Quantization for Medical Image Classification
di: Woergaard, Thomas, et al.
Pubblicazione: (2026)
di: Woergaard, Thomas, et al.
Pubblicazione: (2026)
High-Precision Edge Detection via Task-Adaptive Texture Handling and Ideal-Prior Guidance
di: Shu, Hao
Pubblicazione: (2024)
di: Shu, Hao
Pubblicazione: (2024)
Efficient Edge LLMs Deployment via HessianAware Quantization and CPU GPU Collaborative
di: Zhang, Tuo, et al.
Pubblicazione: (2025)
di: Zhang, Tuo, et al.
Pubblicazione: (2025)
Effective and Efficient Mixed Precision Quantization of Speech Foundation Models
di: Xu, Haoning, et al.
Pubblicazione: (2025)
di: Xu, Haoning, et al.
Pubblicazione: (2025)
MicroMix: Efficient Mixed-Precision Quantization with Microscaling Formats for Large Language Models
di: Liu, Wenyuan, et al.
Pubblicazione: (2025)
di: Liu, Wenyuan, et al.
Pubblicazione: (2025)
Scale When Needed: Adaptive Neuron-level Mixed Precision Quantization Aware Training
di: Varshney, Ayush K., et al.
Pubblicazione: (2026)
di: Varshney, Ayush K., et al.
Pubblicazione: (2026)
KVmix: Gradient-Based Layer Importance-Aware Mixed-Precision Quantization for KV Cache
di: Li, Fei, et al.
Pubblicazione: (2025)
di: Li, Fei, et al.
Pubblicazione: (2025)
Binarization-Aware Adjuster for Discrete Decision Learning with an Application to Edge Detection
di: Shu, Hao
Pubblicazione: (2025)
di: Shu, Hao
Pubblicazione: (2025)
ARQ: A Mixed-Precision Quantization Framework for Accurate and Certifiably Robust DNNs
di: Yang, Yuchen, et al.
Pubblicazione: (2024)
di: Yang, Yuchen, et al.
Pubblicazione: (2024)
Mix-QSAM: Mixed-Precision Quantization of the Segment Anything Model
di: Ranjan, Navin, et al.
Pubblicazione: (2025)
di: Ranjan, Navin, et al.
Pubblicazione: (2025)
Innate immunity in diabetic nephropathy: Pathogenic mechanisms and therapeutic targets
di: Le‐Xin Chen, et al.
Pubblicazione: (2024)
di: Le‐Xin Chen, et al.
Pubblicazione: (2024)
SFMP: Fine-Grained, Hardware-Friendly and Search-Free Mixed-Precision Quantization for Large Language Models
di: Nie, Xin, et al.
Pubblicazione: (2026)
di: Nie, Xin, et al.
Pubblicazione: (2026)
Mix-QViT: Mixed-Precision Vision Transformer Quantization Driven by Layer Importance and Quantization Sensitivity
di: Ranjan, Navin, et al.
Pubblicazione: (2025)
di: Ranjan, Navin, et al.
Pubblicazione: (2025)
GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs
di: Deng, Jianing, et al.
Pubblicazione: (2026)
di: Deng, Jianing, et al.
Pubblicazione: (2026)
MPQ-DMv2: Flexible Residual Mixed Precision Quantization for Low-Bit Diffusion Models with Temporal Distillation
di: Feng, Weilun, et al.
Pubblicazione: (2025)
di: Feng, Weilun, et al.
Pubblicazione: (2025)
DyMoE: Dynamic Expert Orchestration with Mixed-Precision Quantization for Efficient MoE Inference on Edge
di: Huang, Yuegui, et al.
Pubblicazione: (2026)
di: Huang, Yuegui, et al.
Pubblicazione: (2026)
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals
di: Saxena, Utkarsh, et al.
Pubblicazione: (2024)
di: Saxena, Utkarsh, et al.
Pubblicazione: (2024)
Cocktail: Chunk-Adaptive Mixed-Precision Quantization for Long-Context LLM Inference
di: Tao, Wei, et al.
Pubblicazione: (2025)
di: Tao, Wei, et al.
Pubblicazione: (2025)
Occam's Razor and Bender and Koller's Octopus
di: Guerzhoy, Michael
Pubblicazione: (2024)
di: Guerzhoy, Michael
Pubblicazione: (2024)
Mixed-Precision Graph Neural Quantization for Low Bit Large Language Models
di: Liu, Wanlong, et al.
Pubblicazione: (2025)
di: Liu, Wanlong, et al.
Pubblicazione: (2025)
Precisely Skeletal Reorganization via Single Carbon Atom Insertion
di: Bing‐Chao Da, et al.
Pubblicazione: (2026)
di: Bing‐Chao Da, et al.
Pubblicazione: (2026)
Efficient Mixed Precision Quantization in Graph Neural Networks
di: Moustafa, Samir, et al.
Pubblicazione: (2025)
di: Moustafa, Samir, et al.
Pubblicazione: (2025)
MPQ-Diff: Mixed Precision Quantization for Diffusion Models
di: Maruzzelli, Rocco Manz, et al.
Pubblicazione: (2024)
di: Maruzzelli, Rocco Manz, et al.
Pubblicazione: (2024)
Mixed-Precision Quantization for Language Models: Techniques and Prospects
di: Rakka, Mariam, et al.
Pubblicazione: (2025)
di: Rakka, Mariam, et al.
Pubblicazione: (2025)
Documenti analoghi
-
LQA: A Lightweight Quantized-Adaptive Framework for Vision-Language Models on the Edge
di: Wang, Xin, et al.
Pubblicazione: (2026) -
APreQEL: Adaptive Mixed Precision Quantization For Edge LLMs
di: Bouzouad, Meriem, et al.
Pubblicazione: (2026) -
MixKVQ: Query-Aware Mixed-Precision KV Cache Quantization for Long-Context Reasoning
di: Zhang, Tao, et al.
Pubblicazione: (2025) -
AMED: Automatic Mixed-Precision Quantization for Edge Devices
di: Kimhi, Moshe, et al.
Pubblicazione: (2022) -
Generalization Bounds of Spiking Neural Networks via Rademacher Complexity
di: Zhang, Shao-Qun, et al.
Pubblicazione: (2026)