SliM-LLM: Salience-Driven Mixed-Precision Quantization for Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Wei, Qin, Haotong, Liu, Yangdong, Li, Yawei, Liu, Qinshuo, Liu, Xianglong, Benini, Luca, Magno, Michele, Zhang, Shiming, Qi, Xiaojuan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BiLLM: Pushing the Limit of Post-Training Quantization for LLMs
von: Huang, Wei, et al.
Veröffentlicht: (2024)
von: Huang, Wei, et al.
Veröffentlicht: (2024)
On-Chip Hardware-Aware Quantization for Mixed Precision Neural Networks
von: Huang, Wei, et al.
Veröffentlicht: (2023)
von: Huang, Wei, et al.
Veröffentlicht: (2023)
Q-MambaIR: Accurate Quantized Mamba for Efficient Image Restoration
von: Chen, Yujie, et al.
Veröffentlicht: (2025)
von: Chen, Yujie, et al.
Veröffentlicht: (2025)
First-Order Error Matters: Accurate Compensation for Quantized Large Language Models
von: Zheng, Xingyu, et al.
Veröffentlicht: (2025)
von: Zheng, Xingyu, et al.
Veröffentlicht: (2025)
QuantSR+: Pushing the Limit of Quantized Image Super-Resolution Networks
von: Qin, Haotong, et al.
Veröffentlicht: (2026)
von: Qin, Haotong, et al.
Veröffentlicht: (2026)
Enhancing Autonomous Driving Systems with On-Board Deployed Large Language Models
von: Baumann, Nicolas, et al.
Veröffentlicht: (2025)
von: Baumann, Nicolas, et al.
Veröffentlicht: (2025)
Accurate LoRA-Finetuning Quantization of LLMs via Information Retention
von: Qin, Haotong, et al.
Veröffentlicht: (2024)
von: Qin, Haotong, et al.
Veröffentlicht: (2024)
MPQ-DM: Mixed Precision Quantization for Extremely Low Bit Diffusion Models
von: Feng, Weilun, et al.
Veröffentlicht: (2024)
von: Feng, Weilun, et al.
Veröffentlicht: (2024)
Quant-dLLM: Post-Training Extreme Low-Bit Quantization for Diffusion Large Language Models
von: Zhang, Tianao, et al.
Veröffentlicht: (2025)
von: Zhang, Tianao, et al.
Veröffentlicht: (2025)
Event-Priori-Based Vision-Language Model for Efficient Visual Understanding
von: Qin, Haotong, et al.
Veröffentlicht: (2025)
von: Qin, Haotong, et al.
Veröffentlicht: (2025)
An Empirical Study of Qwen3 Quantization
von: Zheng, Xingyu, et al.
Veröffentlicht: (2025)
von: Zheng, Xingyu, et al.
Veröffentlicht: (2025)
Optimal Brain Restoration for Joint Quantization and Sparsification of LLMs
von: Guo, Hang, et al.
Veröffentlicht: (2025)
von: Guo, Hang, et al.
Veröffentlicht: (2025)
An empirical study of LLaMA3 quantization: from LLMs to MLLMs
von: Huang, Wei, et al.
Veröffentlicht: (2024)
von: Huang, Wei, et al.
Veröffentlicht: (2024)
Revisiting Adaptive Rounding with Vectorized Reparameterization for LLM Quantization
von: Zhou, Yuli, et al.
Veröffentlicht: (2026)
von: Zhou, Yuli, et al.
Veröffentlicht: (2026)
BiVM: Accurate Binarized Neural Network for Efficient Video Matting
von: Qin, Haotong, et al.
Veröffentlicht: (2025)
von: Qin, Haotong, et al.
Veröffentlicht: (2025)
A Survey of Low-bit Large Language Models: Basics, Systems, and Algorithms
von: Gong, Ruihao, et al.
Veröffentlicht: (2024)
von: Gong, Ruihao, et al.
Veröffentlicht: (2024)
Post-Training Quantization for Video Matting
von: Zhu, Tianrui, et al.
Veröffentlicht: (2025)
von: Zhu, Tianrui, et al.
Veröffentlicht: (2025)
EdgeCodec: Onboard Lightweight High Fidelity Neural Compressor with Residual Vector Quantization
von: Hodo, Benjamin, et al.
Veröffentlicht: (2025)
von: Hodo, Benjamin, et al.
Veröffentlicht: (2025)
From Fake Focus to Real Precision: Confusion-Driven Adversarial Attention Learning in Transformers
von: Liu, Yawei
Veröffentlicht: (2025)
von: Liu, Yawei
Veröffentlicht: (2025)
Low Latency Visual Inertial Odometry with On-Sensor Accelerated Optical Flow for Resource-Constrained UAVs
von: Kühne, Jonas, et al.
Veröffentlicht: (2024)
von: Kühne, Jonas, et al.
Veröffentlicht: (2024)
RDKV: Rate-Distortion Bit Allocation for Joint Eviction and Quantization of the KV Cache
von: Zhang, Junkai, et al.
Veröffentlicht: (2026)
von: Zhang, Junkai, et al.
Veröffentlicht: (2026)
MPQ-DMv2: Flexible Residual Mixed Precision Quantization for Low-Bit Diffusion Models with Temporal Distillation
von: Feng, Weilun, et al.
Veröffentlicht: (2025)
von: Feng, Weilun, et al.
Veröffentlicht: (2025)
Q-SAM2: Accurate Quantization for Segment Anything Model 2
von: Farronato, Nicola, et al.
Veröffentlicht: (2025)
von: Farronato, Nicola, et al.
Veröffentlicht: (2025)
RobotxR1: Enabling Embodied Robotic Intelligence on Large Language Models through Closed-Loop Reinforcement Learning
von: Boyle, Liam, et al.
Veröffentlicht: (2025)
von: Boyle, Liam, et al.
Veröffentlicht: (2025)
Flexible and Fully Quantized Ultra-Lightweight TinyissimoYOLO for Ultra-Low-Power Edge Systems
von: Moosmann, Julian, et al.
Veröffentlicht: (2023)
von: Moosmann, Julian, et al.
Veröffentlicht: (2023)
Intrinsic Structure as a Proxy for Saliency: SVD-Based Weight Preservation for Mixed-Precision Quantization in Large Language Models
von: Landge, Shashank, et al.
Veröffentlicht: (2025)
von: Landge, Shashank, et al.
Veröffentlicht: (2025)
GWQ: Gradient-Aware Weight Quantization for Large Language Models
von: Shao, Yihua, et al.
Veröffentlicht: (2024)
von: Shao, Yihua, et al.
Veröffentlicht: (2024)
MicroMix: Efficient Mixed-Precision Quantization with Microscaling Formats for Large Language Models
von: Liu, Wenyuan, et al.
Veröffentlicht: (2025)
von: Liu, Wenyuan, et al.
Veröffentlicht: (2025)
Ultra-Efficient On-Device Object Detection on AI-Integrated Smart Glasses with TinyissimoYOLO
von: Moosmann, Julian, et al.
Veröffentlicht: (2023)
von: Moosmann, Julian, et al.
Veröffentlicht: (2023)
BiDM: Pushing the Limit of Quantization for Diffusion Models
von: Zheng, Xingyu, et al.
Veröffentlicht: (2024)
von: Zheng, Xingyu, et al.
Veröffentlicht: (2024)
Saliency-Aware Regularized Quantization Calibration for Large Language Models
von: Zhao, Yanlong, et al.
Veröffentlicht: (2026)
von: Zhao, Yanlong, et al.
Veröffentlicht: (2026)
Mixed-Precision Graph Neural Quantization for Low Bit Large Language Models
von: Liu, Wanlong, et al.
Veröffentlicht: (2025)
von: Liu, Wanlong, et al.
Veröffentlicht: (2025)
Sub-Millisecond Event-Based Eye Tracking on a Resource-Constrained Microcontroller
von: Giordano, Marco, et al.
Veröffentlicht: (2025)
von: Giordano, Marco, et al.
Veröffentlicht: (2025)
Ultra-Lightweight Collaborative Mapping for Robot Swarms
von: Niculescu, Vlad, et al.
Veröffentlicht: (2024)
von: Niculescu, Vlad, et al.
Veröffentlicht: (2024)
NanoSLAM: Enabling Fully Onboard SLAM for Tiny Robots
von: Niculescu, Vlad, et al.
Veröffentlicht: (2023)
von: Niculescu, Vlad, et al.
Veröffentlicht: (2023)
Efficient and Accurate Downfacing Visual Inertial Odometry
von: Kühne, Jonas, et al.
Veröffentlicht: (2025)
von: Kühne, Jonas, et al.
Veröffentlicht: (2025)
LEVIO: Lightweight Embedded Visual Inertial Odometry for Resource-Constrained Devices
von: Kühne, Jonas, et al.
Veröffentlicht: (2026)
von: Kühne, Jonas, et al.
Veröffentlicht: (2026)
BatDeck: Advancing Nano-drone Navigation with Low-power Ultrasound-based Obstacle Avoidance
von: Müller, Hanna, et al.
Veröffentlicht: (2024)
von: Müller, Hanna, et al.
Veröffentlicht: (2024)
Fully Onboard Low-Power Localization with Semantic Sensor Fusion on a Nano-UAV using Floor Plans
von: Zimmerman, Nicky, et al.
Veröffentlicht: (2023)
von: Zimmerman, Nicky, et al.
Veröffentlicht: (2023)
BatDeck -- Ultra Low-power Ultrasonic Ego-velocity Estimation and Obstacle Avoidance on Nano-drones
von: Müller, Hanna, et al.
Veröffentlicht: (2024)
von: Müller, Hanna, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
BiLLM: Pushing the Limit of Post-Training Quantization for LLMs
von: Huang, Wei, et al.
Veröffentlicht: (2024) -
On-Chip Hardware-Aware Quantization for Mixed Precision Neural Networks
von: Huang, Wei, et al.
Veröffentlicht: (2023) -
Q-MambaIR: Accurate Quantized Mamba for Efficient Image Restoration
von: Chen, Yujie, et al.
Veröffentlicht: (2025) -
First-Order Error Matters: Accurate Compensation for Quantized Large Language Models
von: Zheng, Xingyu, et al.
Veröffentlicht: (2025) -
QuantSR+: Pushing the Limit of Quantized Image Super-Resolution Networks
von: Qin, Haotong, et al.
Veröffentlicht: (2026)