Mixed-Precision Quantization for Language Models: Techniques and Prospects
Fuente:
arXiv
Saved in:
| Main Authors: | Rakka, Mariam, Fournarakis, Marios, Krestinskaya, Olga, Bazzi, Jinane, Salama, Khaled N., Kurdahi, Fadi, Eltawil, Ahmed M., Fouda, Mohammed E. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EdgeCIM: A Hardware-Software Co-Design for CIM-Based Acceleration of Small Language Models
by: Bazzi, Jinane, et al.
Published: (2026)
by: Bazzi, Jinane, et al.
Published: (2026)
Joint Hardware-Workload Co-Optimization for In-Memory Computing Accelerators
by: Krestinskaya, Olga, et al.
Published: (2026)
by: Krestinskaya, Olga, et al.
Published: (2026)
CIMNAS: A Joint Framework for Compute-In-Memory-Aware Neural Architecture Search
by: Krestinskaya, Olga, et al.
Published: (2025)
by: Krestinskaya, Olga, et al.
Published: (2025)
Towards Efficient IMC Accelerator Design Through Joint Hardware-Workload Co-optimization
by: Krestinskaya, Olga, et al.
Published: (2024)
by: Krestinskaya, Olga, et al.
Published: (2024)
BF-IMNA: A Bit Fluid In-Memory Neural Architecture for Neural Network Acceleration
by: Rakka, Mariam, et al.
Published: (2024)
by: Rakka, Mariam, et al.
Published: (2024)
SoftmAP: Software-Hardware Co-design for Integer-Only Softmax on Associative Processors
by: Rakka, Mariam, et al.
Published: (2024)
by: Rakka, Mariam, et al.
Published: (2024)
On Jailbreaking Quantized Language Models Through Fault Injection Attacks
by: Zahran, Noureldin, et al.
Published: (2025)
by: Zahran, Noureldin, et al.
Published: (2025)
Sparsity-Aware Streaming SNN Accelerator with Output-Channel Dataflow for Automatic Modulation Classification
by: Yang, Kuilian, et al.
Published: (2026)
by: Yang, Kuilian, et al.
Published: (2026)
A Recurrent YOLOv8-based framework for Event-Based Object Detection
by: Silva, Diego A., et al.
Published: (2024)
by: Silva, Diego A., et al.
Published: (2024)
Chimera: A Block-Based Neural Architecture Search Framework for Event-Based Object Detection
by: Silva, Diego A., et al.
Published: (2024)
by: Silva, Diego A., et al.
Published: (2024)
ConfLayers: Adaptive Confidence-based Layer Skipping for Self-Speculative Decoding
by: Amer, Walaa, et al.
Published: (2026)
by: Amer, Walaa, et al.
Published: (2026)
SalamahBench: Toward Standardized Safety Evaluation for Arabic Language Models
by: Abdelnasser, Omar, et al.
Published: (2026)
by: Abdelnasser, Omar, et al.
Published: (2026)
A Comprehensive Evaluation of Large Language Models on Mental Illnesses in Arabic Context
by: Zahran, Noureldin, et al.
Published: (2025)
by: Zahran, Noureldin, et al.
Published: (2025)
LM-Fix: Lightweight Bit-Flip Detection and Rapid Recovery Framework for Language Models
by: Tahmasivand, Ahmad, et al.
Published: (2025)
by: Tahmasivand, Ahmad, et al.
Published: (2025)
CPTQuant - A Novel Mixed Precision Post-Training Quantization Techniques for Large Language Models
by: Nanda, Amitash, et al.
Published: (2024)
by: Nanda, Amitash, et al.
Published: (2024)
Leveraging Embedding Techniques in Multimodal Machine Learning for Mental Illness Assessment
by: Hassan, Abdelrahaman A., et al.
Published: (2025)
by: Hassan, Abdelrahaman A., et al.
Published: (2025)
Channel-Wise Mixed-Precision Quantization for Large Language Models
by: Chen, Zihan, et al.
Published: (2024)
by: Chen, Zihan, et al.
Published: (2024)
MicroMix: Efficient Mixed-Precision Quantization with Microscaling Formats for Large Language Models
by: Liu, Wenyuan, et al.
Published: (2025)
by: Liu, Wenyuan, et al.
Published: (2025)
Automated Multi-Label Annotation for Mental Health Illnesses Using Large Language Models
by: Hassan, Abdelrahaman A., et al.
Published: (2024)
by: Hassan, Abdelrahaman A., et al.
Published: (2024)
A Comprehensive Evaluation of Large Language Models on Mental Illnesses
by: Hanafi, Abdelrahman, et al.
Published: (2024)
by: Hanafi, Abdelrahman, et al.
Published: (2024)
Mix-QSAM: Mixed-Precision Quantization of the Segment Anything Model
by: Ranjan, Navin, et al.
Published: (2025)
by: Ranjan, Navin, et al.
Published: (2025)
ADDT -- A Digital Twin Framework for Proactive Safety Validation in Autonomous Driving Systems
by: Yu, Bo, et al.
Published: (2025)
by: Yu, Bo, et al.
Published: (2025)
Diagnostic Challenges and Outcome of Classical Phenylketonuria in a Resource‐Constrained Middle Eastern Country
by: Nadine Yazbeck, et al.
Published: (2025)
by: Nadine Yazbeck, et al.
Published: (2025)
Mixed-Precision Graph Neural Quantization for Low Bit Large Language Models
by: Liu, Wanlong, et al.
Published: (2025)
by: Liu, Wanlong, et al.
Published: (2025)
Efficient Mixed Precision Quantization in Graph Neural Networks
by: Moustafa, Samir, et al.
Published: (2025)
by: Moustafa, Samir, et al.
Published: (2025)
Leveraging Audio and Text Modalities in Mental Health: A Study of LLMs Performance
by: Ali, Abdelrahman A., et al.
Published: (2024)
by: Ali, Abdelrahman A., et al.
Published: (2024)
PsychiatryBench: A Multi-Task Benchmark for LLMs in Psychiatry
by: Fouda, Aya E., et al.
Published: (2025)
by: Fouda, Aya E., et al.
Published: (2025)
MPQ-Diff: Mixed Precision Quantization for Diffusion Models
by: Maruzzelli, Rocco Manz, et al.
Published: (2024)
by: Maruzzelli, Rocco Manz, et al.
Published: (2024)
OMPQ: Orthogonal Mixed Precision Quantization
by: Ma, Yuexiao, et al.
Published: (2021)
by: Ma, Yuexiao, et al.
Published: (2021)
SliM-LLM: Salience-Driven Mixed-Precision Quantization for Large Language Models
by: Huang, Wei, et al.
Published: (2024)
by: Huang, Wei, et al.
Published: (2024)
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals
by: Saxena, Utkarsh, et al.
Published: (2024)
by: Saxena, Utkarsh, et al.
Published: (2024)
APTQ: Attention-aware Post-Training Mixed-Precision Quantization for Large Language Models
by: Guan, Ziyi, et al.
Published: (2024)
by: Guan, Ziyi, et al.
Published: (2024)
At the Dawn of Generative AI Era: A Tutorial-cum-Survey on New Frontiers in 6G Wireless Intelligence
by: Celik, Abdulkadir, et al.
Published: (2024)
by: Celik, Abdulkadir, et al.
Published: (2024)
Effective and Efficient Mixed Precision Quantization of Speech Foundation Models
by: Xu, Haoning, et al.
Published: (2025)
by: Xu, Haoning, et al.
Published: (2025)
Nivel de conocimiento sobre uso profiláctico y terapéutico del zinc en edades pediátricas
by: Olga María Blanco Bazzi
Published: (2013)
by: Olga María Blanco Bazzi
Published: (2013)
Algunos aspectos relacionados con el zinc como elemento esencial en la nutrición infantil
by: Olga María Blanco Bazzi
Published: (2013)
by: Olga María Blanco Bazzi
Published: (2013)
INTERVENCION TEMPRANA EN NIÑOS CON TRASTORNOS DEL NEURODESARROLLO.
by: Olga María Blanco Bazzi
Published: (2005)
by: Olga María Blanco Bazzi
Published: (2005)
Assessing Barriers to the Adoption of Green Technologies in Developing Countries by Using an Integrated Decision‐Making System
by: Muhammad Ikram, et al.
Published: (2025)
by: Muhammad Ikram, et al.
Published: (2025)
Mix-QViT: Mixed-Precision Vision Transformer Quantization Driven by Layer Importance and Quantization Sensitivity
by: Ranjan, Navin, et al.
Published: (2025)
by: Ranjan, Navin, et al.
Published: (2025)
Mixed-Precision Quantization for Deep Vision Models with Integer Quadratic Programming
by: Deng, Zihao, et al.
Published: (2023)
by: Deng, Zihao, et al.
Published: (2023)
Similar Items
-
EdgeCIM: A Hardware-Software Co-Design for CIM-Based Acceleration of Small Language Models
by: Bazzi, Jinane, et al.
Published: (2026) -
Joint Hardware-Workload Co-Optimization for In-Memory Computing Accelerators
by: Krestinskaya, Olga, et al.
Published: (2026) -
CIMNAS: A Joint Framework for Compute-In-Memory-Aware Neural Architecture Search
by: Krestinskaya, Olga, et al.
Published: (2025) -
Towards Efficient IMC Accelerator Design Through Joint Hardware-Workload Co-optimization
by: Krestinskaya, Olga, et al.
Published: (2024) -
BF-IMNA: A Bit Fluid In-Memory Neural Architecture for Neural Network Acceleration
by: Rakka, Mariam, et al.
Published: (2024)