InfoQ: Mixed-Precision Quantization via Global Information Flow
Fuente:
arXiv
Saved in:
| Main Authors: | Akbulut, Mehmet Emre, Shalby, Hazem Hesham Yousef, Pittorino, Fabrizio, Roveri, Manuel |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DQT: Dynamic Quantization Training via Dequantization-Free Nested Integer Arithmetic
by: Shalby, Hazem Hesham Yousef, et al.
Published: (2025)
by: Shalby, Hazem Hesham Yousef, et al.
Published: (2025)
Dendron: Enhancing Human Activity Recognition with On-Device TinyML Learning
by: Shalby, Hazem Hesham Yousef, et al.
Published: (2025)
by: Shalby, Hazem Hesham Yousef, et al.
Published: (2025)
EmbBERT: Attention Under 2 MB Memory
by: Bravin, Riccardo, et al.
Published: (2025)
by: Bravin, Riccardo, et al.
Published: (2025)
Position Paper: From Edge AI to Adaptive Edge AI
by: Pittorino, Fabrizio, et al.
Published: (2026)
by: Pittorino, Fabrizio, et al.
Published: (2026)
StreamTinyNet: video streaming analysis with spatial-temporal TinyML
by: Shalby, Hazem Hesham Yousef, et al.
Published: (2024)
by: Shalby, Hazem Hesham Yousef, et al.
Published: (2024)
Quantifying Cryptocurrency Unpredictability: A Comprehensive Study of Complexity and Forecasting
by: Puoti, Francesco, et al.
Published: (2025)
by: Puoti, Francesco, et al.
Published: (2025)
On-Sensor Convolutional Neural Networks with Early-Exits
by: Shalby, Hazem Hesham Yousef, et al.
Published: (2025)
by: Shalby, Hazem Hesham Yousef, et al.
Published: (2025)
FlatNAS: optimizing Flatness in Neural Architecture Search for Out-of-Distribution Robustness
by: Gambella, Matteo, et al.
Published: (2024)
by: Gambella, Matteo, et al.
Published: (2024)
Architecture-Aware Minimization (A$^2$M): How to Find Flat Minima in Neural Architecture Search
by: Gambella, Matteo, et al.
Published: (2025)
by: Gambella, Matteo, et al.
Published: (2025)
An Algorithm for On-Sensor Agnostic Detection of Changes in Human Activity for Ultra-Low-Power Applications
by: Rimoldi, Sara, et al.
Published: (2026)
by: Rimoldi, Sara, et al.
Published: (2026)
Training Multi-Layer Binary Neural Networks With Local Binary Error Signals
by: Colombo, Luca, et al.
Published: (2024)
by: Colombo, Luca, et al.
Published: (2024)
HERCULES: Hardware-Efficient, Robust, Continual Learning Neural Architecture Search
by: Gambella, Matteo, et al.
Published: (2026)
by: Gambella, Matteo, et al.
Published: (2026)
SQUAD: Scalable Quorum Adaptive Decisions via ensemble of early exit neural networks
by: Gambella, Matteo, et al.
Published: (2026)
by: Gambella, Matteo, et al.
Published: (2026)
What changes after deployment? A survey on On-device Learning in TinyML
by: Pavan, Massimo, et al.
Published: (2026)
by: Pavan, Massimo, et al.
Published: (2026)
Active In-Context Learning for Tabular Foundation Models
by: Treerath, Wilailuck, et al.
Published: (2026)
by: Treerath, Wilailuck, et al.
Published: (2026)
BEP: A Binary Error Propagation Algorithm for Binary Neural Networks Training
by: Colombo, Luca, et al.
Published: (2025)
by: Colombo, Luca, et al.
Published: (2025)
CoopQ: Cooperative Game Inspired Layerwise Mixed Precision Quantization for LLMs
by: Zhao, Junchen, et al.
Published: (2025)
by: Zhao, Junchen, et al.
Published: (2025)
DiRotQ: Rotation-Aware Quantization for 4-bit Diffusion Transformers
by: Sharify, Sayeh, et al.
Published: (2026)
by: Sharify, Sayeh, et al.
Published: (2026)
OMPQ: Orthogonal Mixed Precision Quantization
by: Ma, Yuexiao, et al.
Published: (2021)
by: Ma, Yuexiao, et al.
Published: (2021)
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals
by: Saxena, Utkarsh, et al.
Published: (2024)
by: Saxena, Utkarsh, et al.
Published: (2024)
FlexiQ: Adaptive Mixed-Precision Quantization for Latency/Accuracy Trade-Offs in Deep Neural Networks
by: Kim, Jaemin, et al.
Published: (2025)
by: Kim, Jaemin, et al.
Published: (2025)
GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs
by: Deng, Jianing, et al.
Published: (2026)
by: Deng, Jianing, et al.
Published: (2026)
InfoFlow KV: Information-Flow-Aware KV Recomputation for Long Context
by: Teng, Xin, et al.
Published: (2026)
by: Teng, Xin, et al.
Published: (2026)
Efficient Mixed Precision Quantization in Graph Neural Networks
by: Moustafa, Samir, et al.
Published: (2025)
by: Moustafa, Samir, et al.
Published: (2025)
MoPEQ: Mixture of Mixed Precision Quantized Experts
by: Chitty-Venkata, Krishna Teja, et al.
Published: (2025)
by: Chitty-Venkata, Krishna Teja, et al.
Published: (2025)
AMED: Automatic Mixed-Precision Quantization for Edge Devices
by: Kimhi, Moshe, et al.
Published: (2022)
by: Kimhi, Moshe, et al.
Published: (2022)
FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs
by: Xie, Xilong, et al.
Published: (2025)
by: Xie, Xilong, et al.
Published: (2025)
Mixed-Precision Quantization for Language Models: Techniques and Prospects
by: Rakka, Mariam, et al.
Published: (2025)
by: Rakka, Mariam, et al.
Published: (2025)
TActiLE: Tiny Active LEarning for wearable devices
by: Pavan, Massimo, et al.
Published: (2025)
by: Pavan, Massimo, et al.
Published: (2025)
STaMP: Sequence Transformation and Mixed Precision for Low-Precision Activation Quantization
by: Federici, Marco, et al.
Published: (2025)
by: Federici, Marco, et al.
Published: (2025)
InfoBridge: Mutual Information estimation via Bridge Matching
by: Kholkin, Sergei, et al.
Published: (2025)
by: Kholkin, Sergei, et al.
Published: (2025)
Flexible Mixed Precision Quantization for Learned Image Compression
by: Hossain, Md Adnan Faisal, et al.
Published: (2025)
by: Hossain, Md Adnan Faisal, et al.
Published: (2025)
APreQEL: Adaptive Mixed Precision Quantization For Edge LLMs
by: Bouzouad, Meriem, et al.
Published: (2026)
by: Bouzouad, Meriem, et al.
Published: (2026)
Where and How to Enhance: Discovering Bit-Width Contribution for Mixed Precision Quantization
by: Kang, Haidong, et al.
Published: (2025)
by: Kang, Haidong, et al.
Published: (2025)
LoRAQuant: Mixed-Precision Quantization of LoRA to Ultra-Low Bits
by: Mirzaei, Amir Reza, et al.
Published: (2025)
by: Mirzaei, Amir Reza, et al.
Published: (2025)
Data-Free Quantization via Mixed-Precision Compensation without Fine-Tuning
by: Chen, Jun, et al.
Published: (2023)
by: Chen, Jun, et al.
Published: (2023)
MixDiT: Accelerating Image Diffusion Transformer Inference with Mixed-Precision MX Quantization
by: Kim, Daeun, et al.
Published: (2025)
by: Kim, Daeun, et al.
Published: (2025)
MicroMix: Efficient Mixed-Precision Quantization with Microscaling Formats for Large Language Models
by: Liu, Wenyuan, et al.
Published: (2025)
by: Liu, Wenyuan, et al.
Published: (2025)
InfoQuant: Shaping Activation Distributions for Low-Bit LLM Quantization
by: Li, Ke, et al.
Published: (2026)
by: Li, Ke, et al.
Published: (2026)
MPQ-Diff: Mixed Precision Quantization for Diffusion Models
by: Maruzzelli, Rocco Manz, et al.
Published: (2024)
by: Maruzzelli, Rocco Manz, et al.
Published: (2024)
Similar Items
-
DQT: Dynamic Quantization Training via Dequantization-Free Nested Integer Arithmetic
by: Shalby, Hazem Hesham Yousef, et al.
Published: (2025) -
Dendron: Enhancing Human Activity Recognition with On-Device TinyML Learning
by: Shalby, Hazem Hesham Yousef, et al.
Published: (2025) -
EmbBERT: Attention Under 2 MB Memory
by: Bravin, Riccardo, et al.
Published: (2025) -
Position Paper: From Edge AI to Adaptive Edge AI
by: Pittorino, Fabrizio, et al.
Published: (2026) -
StreamTinyNet: video streaming analysis with spatial-temporal TinyML
by: Shalby, Hazem Hesham Yousef, et al.
Published: (2024)