Scale When Needed: Adaptive Neuron-level Mixed Precision Quantization Aware Training
Fuente:
arXiv
Saved in:
| Main Authors: | Varshney, Ayush K., Vandikas, Konstantinos, Girdzijauskas, Šarūnas, Orucu, Adam, Feljan, Aneta Vulgarakis |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When to restart? Exploring escalating restarts on convergence
by: Varshney, Ayush K., et al.
Published: (2026)
by: Varshney, Ayush K., et al.
Published: (2026)
Unlearning Clients, Features and Samples in Vertical Federated Learning
by: Varshney, Ayush K., et al.
Published: (2025)
by: Varshney, Ayush K., et al.
Published: (2025)
Multi-agent transformer-accelerated RL for satisfaction of STL specifications
by: Forsberg, Albin Larsson, et al.
Published: (2024)
by: Forsberg, Albin Larsson, et al.
Published: (2024)
Reinforcement Learning for Long-Horizon Unordered Tasks: From Boolean to Coupled Reward Machines
by: Levina, Kristina, et al.
Published: (2025)
by: Levina, Kristina, et al.
Published: (2025)
Reinforcement Learning with Reward Machines for Sleep Control in Mobile Networks
by: Levina, Kristina, et al.
Published: (2026)
by: Levina, Kristina, et al.
Published: (2026)
Numeric Reward Machines
by: Levina, Kristina, et al.
Published: (2024)
by: Levina, Kristina, et al.
Published: (2024)
Automated Model Design using Gated Neuron Selection in Telecom
by: Orucu, Adam, et al.
Published: (2026)
by: Orucu, Adam, et al.
Published: (2026)
A Survey on the Integration of Generative AI for Critical Thinking in Mobile Networks
by: Karapantelakis, Athanasios, et al.
Published: (2024)
by: Karapantelakis, Athanasios, et al.
Published: (2024)
Are We Wasting Time? A Fast, Accurate Performance Evaluation Framework for Knowledge Graph Link Predictors
by: Cornell, Filip, et al.
Published: (2024)
by: Cornell, Filip, et al.
Published: (2024)
Open World Learning Graph Convolution for Latency Estimation in Routing Networks
by: Jin, Yifei, et al.
Published: (2022)
by: Jin, Yifei, et al.
Published: (2022)
Semi-decentralized Training of Spatio-Temporal Graph Neural Networks for Traffic Prediction
by: Kralj, Ivan, et al.
Published: (2024)
by: Kralj, Ivan, et al.
Published: (2024)
Adaptive Graph Pruning with Sudden-Events Evaluation for Traffic Prediction using Online Semi-Decentralized ST-GNNs
by: Kralj, Ivan, et al.
Published: (2025)
by: Kralj, Ivan, et al.
Published: (2025)
SANDWICH: Towards an Offline, Differentiable, Fully-Trainable Wireless Neural Ray-Tracing Surrogate
by: Jin, Yifei, et al.
Published: (2024)
by: Jin, Yifei, et al.
Published: (2024)
APreQEL: Adaptive Mixed Precision Quantization For Edge LLMs
by: Bouzouad, Meriem, et al.
Published: (2026)
by: Bouzouad, Meriem, et al.
Published: (2026)
On-Chip Hardware-Aware Quantization for Mixed Precision Neural Networks
by: Huang, Wei, et al.
Published: (2023)
by: Huang, Wei, et al.
Published: (2023)
MixKVQ: Query-Aware Mixed-Precision KV Cache Quantization for Long-Context Reasoning
by: Zhang, Tao, et al.
Published: (2025)
by: Zhang, Tao, et al.
Published: (2025)
RAMP: Reinforcement Adaptive Mixed Precision Quantization for Efficient On Device LLM Inference
by: Gautam, Arpit Singh, et al.
Published: (2026)
by: Gautam, Arpit Singh, et al.
Published: (2026)
Towards Neural Architecture Search for Transfer Learning in 6G Networks
by: Orucu, Adam, et al.
Published: (2024)
by: Orucu, Adam, et al.
Published: (2024)
Learning from Loss Landscape: Generalizable Mixed-Precision Quantization via Adaptive Sharpness-Aware Gradient Aligning
by: Ma, Lianbo, et al.
Published: (2025)
by: Ma, Lianbo, et al.
Published: (2025)
AdaQAT: Adaptive Bit-Width Quantization-Aware Training
by: Gernigon, Cédric, et al.
Published: (2024)
by: Gernigon, Cédric, et al.
Published: (2024)
Mixed-Precision Quantization for Language Models: Techniques and Prospects
by: Rakka, Mariam, et al.
Published: (2025)
by: Rakka, Mariam, et al.
Published: (2025)
APTQ: Attention-aware Post-Training Mixed-Precision Quantization for Large Language Models
by: Guan, Ziyi, et al.
Published: (2024)
by: Guan, Ziyi, et al.
Published: (2024)
STaMP: Sequence Transformation and Mixed Precision for Low-Precision Activation Quantization
by: Federici, Marco, et al.
Published: (2025)
by: Federici, Marco, et al.
Published: (2025)
No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization
by: Yang, June Yong, et al.
Published: (2024)
by: Yang, June Yong, et al.
Published: (2024)
Channel-Wise Mixed-Precision Quantization for Large Language Models
by: Chen, Zihan, et al.
Published: (2024)
by: Chen, Zihan, et al.
Published: (2024)
EdgeRazor: A Lightweight Framework for Large Language Models via Mixed-Precision Quantization-Aware Distillation
by: Zhang, Shu-Hao, et al.
Published: (2026)
by: Zhang, Shu-Hao, et al.
Published: (2026)
SASQ: Static Activation Scaling for Quantization-Aware Training in Large Language Models
by: Mao, Shizhuo, et al.
Published: (2025)
by: Mao, Shizhuo, et al.
Published: (2025)
MicroMix: Efficient Mixed-Precision Quantization with Microscaling Formats for Large Language Models
by: Liu, Wenyuan, et al.
Published: (2025)
by: Liu, Wenyuan, et al.
Published: (2025)
AutoScale: Scale-Aware Data Mixing for Pre-Training LLMs
by: Kang, Feiyang, et al.
Published: (2024)
by: Kang, Feiyang, et al.
Published: (2024)
Compute-Optimal Quantization-Aware Training
by: Dremov, Aleksandr, et al.
Published: (2025)
by: Dremov, Aleksandr, et al.
Published: (2025)
How to Parameterize Asymmetric Quantization Ranges for Quantization-Aware Training
by: You, Jaeseong, et al.
Published: (2024)
by: You, Jaeseong, et al.
Published: (2024)
GAC: Noise-Aware Adaptive Mixing for Hybrid SFT-RL Post-Training
by: Hu, Yuelin, et al.
Published: (2026)
by: Hu, Yuelin, et al.
Published: (2026)
Effective and Efficient Mixed Precision Quantization of Speech Foundation Models
by: Xu, Haoning, et al.
Published: (2025)
by: Xu, Haoning, et al.
Published: (2025)
Do Language Models Know When They're Hallucinating References?
by: Agrawal, Ayush, et al.
Published: (2023)
by: Agrawal, Ayush, et al.
Published: (2023)
KVmix: Gradient-Based Layer Importance-Aware Mixed-Precision Quantization for KV Cache
by: Li, Fei, et al.
Published: (2025)
by: Li, Fei, et al.
Published: (2025)
A Metric Driven Approach to Mixed Precision Training
by: Rasquinha, Mitchelle, et al.
Published: (2024)
by: Rasquinha, Mitchelle, et al.
Published: (2024)
MoR: Mixture Of Representations For Mixed-Precision Training
by: Su, Bor-Yiing, et al.
Published: (2025)
by: Su, Bor-Yiing, et al.
Published: (2025)
KVTuner: Sensitivity-Aware Layer-Wise Mixed-Precision KV Cache Quantization for Efficient and Nearly Lossless LLM Inference
by: Li, Xing, et al.
Published: (2025)
by: Li, Xing, et al.
Published: (2025)
A Quantized VAE-MLP Botnet Detection Model: A Systematic Evaluation of Quantization-Aware Training and Post-Training Quantization Strategies
by: Wasswa, Hassan, et al.
Published: (2025)
by: Wasswa, Hassan, et al.
Published: (2025)
Beyond Pooling: Matching for Robust Generalization under Data Heterogeneity
by: Roy, Ayush, et al.
Published: (2026)
by: Roy, Ayush, et al.
Published: (2026)
Similar Items
-
When to restart? Exploring escalating restarts on convergence
by: Varshney, Ayush K., et al.
Published: (2026) -
Unlearning Clients, Features and Samples in Vertical Federated Learning
by: Varshney, Ayush K., et al.
Published: (2025) -
Multi-agent transformer-accelerated RL for satisfaction of STL specifications
by: Forsberg, Albin Larsson, et al.
Published: (2024) -
Reinforcement Learning for Long-Horizon Unordered Tasks: From Boolean to Coupled Reward Machines
by: Levina, Kristina, et al.
Published: (2025) -
Reinforcement Learning with Reward Machines for Sleep Control in Mobile Networks
by: Levina, Kristina, et al.
Published: (2026)