DQT: Dynamic Quantization Training via Dequantization-Free Nested Integer Arithmetic
Fuente:
arXiv
Guardado en:
| Autores principales: | Shalby, Hazem Hesham Yousef, Pittorino, Fabrizio, Palermo, Francesca, Trojaniello, Diana, Roveri, Manuel |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Dendron: Enhancing Human Activity Recognition with On-Device TinyML Learning
por: Shalby, Hazem Hesham Yousef, et al.
Publicado: (2025)
por: Shalby, Hazem Hesham Yousef, et al.
Publicado: (2025)
InfoQ: Mixed-Precision Quantization via Global Information Flow
por: Akbulut, Mehmet Emre, et al.
Publicado: (2025)
por: Akbulut, Mehmet Emre, et al.
Publicado: (2025)
On-Sensor Convolutional Neural Networks with Early-Exits
por: Shalby, Hazem Hesham Yousef, et al.
Publicado: (2025)
por: Shalby, Hazem Hesham Yousef, et al.
Publicado: (2025)
StreamTinyNet: video streaming analysis with spatial-temporal TinyML
por: Shalby, Hazem Hesham Yousef, et al.
Publicado: (2024)
por: Shalby, Hazem Hesham Yousef, et al.
Publicado: (2024)
EmbBERT: Attention Under 2 MB Memory
por: Bravin, Riccardo, et al.
Publicado: (2025)
por: Bravin, Riccardo, et al.
Publicado: (2025)
Position Paper: From Edge AI to Adaptive Edge AI
por: Pittorino, Fabrizio, et al.
Publicado: (2026)
por: Pittorino, Fabrizio, et al.
Publicado: (2026)
FlatNAS: optimizing Flatness in Neural Architecture Search for Out-of-Distribution Robustness
por: Gambella, Matteo, et al.
Publicado: (2024)
por: Gambella, Matteo, et al.
Publicado: (2024)
An Algorithm for On-Sensor Agnostic Detection of Changes in Human Activity for Ultra-Low-Power Applications
por: Rimoldi, Sara, et al.
Publicado: (2026)
por: Rimoldi, Sara, et al.
Publicado: (2026)
What changes after deployment? A survey on On-device Learning in TinyML
por: Pavan, Massimo, et al.
Publicado: (2026)
por: Pavan, Massimo, et al.
Publicado: (2026)
BEP: A Binary Error Propagation Algorithm for Binary Neural Networks Training
por: Colombo, Luca, et al.
Publicado: (2025)
por: Colombo, Luca, et al.
Publicado: (2025)
NestQuant: Post-Training Integer-Nesting Quantization for On-Device DNN
por: Xie, Jianhang, et al.
Publicado: (2025)
por: Xie, Jianhang, et al.
Publicado: (2025)
Training Multi-Layer Binary Neural Networks With Local Binary Error Signals
por: Colombo, Luca, et al.
Publicado: (2024)
por: Colombo, Luca, et al.
Publicado: (2024)
Quantifying Cryptocurrency Unpredictability: A Comprehensive Study of Complexity and Forecasting
por: Puoti, Francesco, et al.
Publicado: (2025)
por: Puoti, Francesco, et al.
Publicado: (2025)
NITRO-D: Native Integer-only Training of Deep Convolutional Neural Networks
por: Pirillo, Alberto, et al.
Publicado: (2024)
por: Pirillo, Alberto, et al.
Publicado: (2024)
Multi-Scale Dequant: Eliminating Dequantization Bottleneck via Activation Decomposition for Efficient LLM Inference
por: Zheng, Lingchao, et al.
Publicado: (2026)
por: Zheng, Lingchao, et al.
Publicado: (2026)
TIFeD: a Tiny Integer-based Federated learning algorithm with Direct feedback alignment
por: Colombo, Luca, et al.
Publicado: (2024)
por: Colombo, Luca, et al.
Publicado: (2024)
Architecture-Aware Minimization (A$^2$M): How to Find Flat Minima in Neural Architecture Search
por: Gambella, Matteo, et al.
Publicado: (2025)
por: Gambella, Matteo, et al.
Publicado: (2025)
HERCULES: Hardware-Efficient, Robust, Continual Learning Neural Architecture Search
por: Gambella, Matteo, et al.
Publicado: (2026)
por: Gambella, Matteo, et al.
Publicado: (2026)
Federated Reinforcement Learning for Runtime Optimization of AI Applications in Smart Eyewears
por: Sedghani, Hamta, et al.
Publicado: (2025)
por: Sedghani, Hamta, et al.
Publicado: (2025)
SQUAD: Scalable Quorum Adaptive Decisions via ensemble of early exit neural networks
por: Gambella, Matteo, et al.
Publicado: (2026)
por: Gambella, Matteo, et al.
Publicado: (2026)
Arithmetic-Intensity-Aware Quantization
por: Singh, Taig, et al.
Publicado: (2025)
por: Singh, Taig, et al.
Publicado: (2025)
Integer Scale: A Free Lunch for Faster Fine-grained Quantization of LLMs
por: Li, Qingyuan, et al.
Publicado: (2024)
por: Li, Qingyuan, et al.
Publicado: (2024)
Gradient-Free Training of Quantized Neural Networks
por: Cohen, Noa, et al.
Publicado: (2024)
por: Cohen, Noa, et al.
Publicado: (2024)
Learning Quantized Continuous Controllers for Integer Hardware
por: Kresse, Fabian, et al.
Publicado: (2025)
por: Kresse, Fabian, et al.
Publicado: (2025)
NestQuant: Nested Lattice Quantization for Matrix Products and LLMs
por: Savkin, Semyon, et al.
Publicado: (2025)
por: Savkin, Semyon, et al.
Publicado: (2025)
Calibration and Transformation-Free Weight-Only LLMs Quantization via Dynamic Grouping
por: Zheng, Xinzhe, et al.
Publicado: (2025)
por: Zheng, Xinzhe, et al.
Publicado: (2025)
Zero-Shot Quantization via Weight-Space Arithmetic
por: Solombrino, Daniele, et al.
Publicado: (2026)
por: Solombrino, Daniele, et al.
Publicado: (2026)
SEAL: Searching Expandable Architectures for Incremental Learning
por: Gambella, Matteo, et al.
Publicado: (2025)
por: Gambella, Matteo, et al.
Publicado: (2025)
DYNAMAX: Dynamic computing for Transformers and Mamba based architectures
por: Nogales, Miguel, et al.
Publicado: (2025)
por: Nogales, Miguel, et al.
Publicado: (2025)
Task Arithmetic in Trust Region: A Training-Free Model Merging Approach to Navigate Knowledge Conflicts
por: Sun, Wenju, et al.
Publicado: (2025)
por: Sun, Wenju, et al.
Publicado: (2025)
GSQ-Tuning: Group-Shared Exponents Integer in Fully Quantized Training for LLMs On-Device Fine-tuning
por: Zhou, Sifan, et al.
Publicado: (2025)
por: Zhou, Sifan, et al.
Publicado: (2025)
Integer-only Quantized Transformers for Embedded FPGA-based Time-series Forecasting in AIoT
por: Ling, Tianheng, et al.
Publicado: (2024)
por: Ling, Tianheng, et al.
Publicado: (2024)
OptRot: Mitigating Weight Outliers via Data-Free Rotations for Post-Training Quantization
por: Gadhikar, Advait, et al.
Publicado: (2025)
por: Gadhikar, Advait, et al.
Publicado: (2025)
D-QRELO: Training- and Data-Free Delta Compression for Large Language Models via Quantization and Residual Low-Rank Approximation
por: Li, Junlin, et al.
Publicado: (2026)
por: Li, Junlin, et al.
Publicado: (2026)
Quamba: A Post-Training Quantization Recipe for Selective State Space Models
por: Chiang, Hung-Yueh, et al.
Publicado: (2024)
por: Chiang, Hung-Yueh, et al.
Publicado: (2024)
Bitwidth-Specific Logarithmic Arithmetic for Future Hardware-Accelerated Training
por: Hamad, Hassan, et al.
Publicado: (2025)
por: Hamad, Hassan, et al.
Publicado: (2025)
How to Parameterize Asymmetric Quantization Ranges for Quantization-Aware Training
por: You, Jaeseong, et al.
Publicado: (2024)
por: You, Jaeseong, et al.
Publicado: (2024)
Distilling Linearized Behavior into Non-Linear Fine-Tuning for Effective Task Arithmetic
por: Sommariva, Thomas, et al.
Publicado: (2026)
por: Sommariva, Thomas, et al.
Publicado: (2026)
A Quantized VAE-MLP Botnet Detection Model: A Systematic Evaluation of Quantization-Aware Training and Post-Training Quantization Strategies
por: Wasswa, Hassan, et al.
Publicado: (2025)
por: Wasswa, Hassan, et al.
Publicado: (2025)
Shapley-PC: Constraint-based Causal Structure Learning with a Shapley Inspired Framework
por: Russo, Fabrizio, et al.
Publicado: (2023)
por: Russo, Fabrizio, et al.
Publicado: (2023)
Ejemplares similares
-
Dendron: Enhancing Human Activity Recognition with On-Device TinyML Learning
por: Shalby, Hazem Hesham Yousef, et al.
Publicado: (2025) -
InfoQ: Mixed-Precision Quantization via Global Information Flow
por: Akbulut, Mehmet Emre, et al.
Publicado: (2025) -
On-Sensor Convolutional Neural Networks with Early-Exits
por: Shalby, Hazem Hesham Yousef, et al.
Publicado: (2025) -
StreamTinyNet: video streaming analysis with spatial-temporal TinyML
por: Shalby, Hazem Hesham Yousef, et al.
Publicado: (2024) -
EmbBERT: Attention Under 2 MB Memory
por: Bravin, Riccardo, et al.
Publicado: (2025)