Arithmetic-Intensity-Aware Quantization
Fuente:
arXiv
Saved in:
| Main Authors: | Singh, Taig, Rajan, Shreshth, Jain, Nikhil |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MultiVer: Zero-Shot Multi-Agent Vulnerability Detection
by: Rajan, Shreshth
Published: (2026)
by: Rajan, Shreshth
Published: (2026)
TAGMol: Target-Aware Gradient-guided Molecule Generation
by: Dorna, Vineeth, et al.
Published: (2024)
by: Dorna, Vineeth, et al.
Published: (2024)
PeerRank: Autonomous LLM Evaluation Through Web-Grounded, Bias-Controlled Peer Review
by: Margalit, Yanki, et al.
Published: (2026)
by: Margalit, Yanki, et al.
Published: (2026)
DQT: Dynamic Quantization Training via Dequantization-Free Nested Integer Arithmetic
by: Shalby, Hazem Hesham Yousef, et al.
Published: (2025)
by: Shalby, Hazem Hesham Yousef, et al.
Published: (2025)
A Selective Quantization Tuner for ONNX Models
by: Louloudakis, Nikolaos, et al.
Published: (2025)
by: Louloudakis, Nikolaos, et al.
Published: (2025)
Carbon Intensity-Aware Adaptive Inference of DNNs
by: Jung, Jiwan
Published: (2024)
by: Jung, Jiwan
Published: (2024)
Matryoshka Quantization
by: Nair, Pranav, et al.
Published: (2025)
by: Nair, Pranav, et al.
Published: (2025)
TABES: Trajectory-Aware Backward-on-Entropy Steering for Masked Diffusion Models
by: Saini, Shreshth, et al.
Published: (2026)
by: Saini, Shreshth, et al.
Published: (2026)
How to Parameterize Asymmetric Quantization Ranges for Quantization-Aware Training
by: You, Jaeseong, et al.
Published: (2024)
by: You, Jaeseong, et al.
Published: (2024)
Transformers Can Do Arithmetic with the Right Embeddings
by: McLeish, Sean, et al.
Published: (2024)
by: McLeish, Sean, et al.
Published: (2024)
Compute-Optimal Quantization-Aware Training
by: Dremov, Aleksandr, et al.
Published: (2025)
by: Dremov, Aleksandr, et al.
Published: (2025)
Multi-Agent Code Verification via Information Theory
by: Rajan, Shreshth
Published: (2025)
by: Rajan, Shreshth
Published: (2025)
Zero-Shot Quantization via Weight-Space Arithmetic
by: Solombrino, Daniele, et al.
Published: (2026)
by: Solombrino, Daniele, et al.
Published: (2026)
FAQ: Mitigating Quantization Error via Regenerating Calibration Data with Family-Aware Quantization
by: Xiao, Haiyang, et al.
Published: (2026)
by: Xiao, Haiyang, et al.
Published: (2026)
On-the-Fly Adaptation to Quantization: Configuration-Aware LoRA for Efficient Fine-Tuning of Quantized LLMs
by: Ye, Rongguang, et al.
Published: (2025)
by: Ye, Rongguang, et al.
Published: (2025)
A Quantized VAE-MLP Botnet Detection Model: A Systematic Evaluation of Quantization-Aware Training and Post-Training Quantization Strategies
by: Wasswa, Hassan, et al.
Published: (2025)
by: Wasswa, Hassan, et al.
Published: (2025)
QUAIL: Quantization Aware Unlearning for Mitigating Misinformation in LLMs
by: Mishra, Himanshu, et al.
Published: (2026)
by: Mishra, Himanshu, et al.
Published: (2026)
Pioneering 4-Bit FP Quantization for Diffusion Models: Mixup-Sign Quantization and Timestep-Aware Fine-Tuning
by: Zhao, Maosen, et al.
Published: (2025)
by: Zhao, Maosen, et al.
Published: (2025)
Saliency-Aware Regularized Quantization Calibration for Large Language Models
by: Zhao, Yanlong, et al.
Published: (2026)
by: Zhao, Yanlong, et al.
Published: (2026)
AdaQAT: Adaptive Bit-Width Quantization-Aware Training
by: Gernigon, Cédric, et al.
Published: (2024)
by: Gernigon, Cédric, et al.
Published: (2024)
Attn-QAT: 4-Bit Attention With Quantization-Aware Training
by: Zhang, Peiyuan, et al.
Published: (2026)
by: Zhang, Peiyuan, et al.
Published: (2026)
EXAQ: Exponent Aware Quantization For LLMs Acceleration
by: Shkolnik, Moran, et al.
Published: (2024)
by: Shkolnik, Moran, et al.
Published: (2024)
Uncertainty-Gated Region-Level Retrieval for Robust Semantic Segmentation
by: Rajan, Shreshth, et al.
Published: (2025)
by: Rajan, Shreshth, et al.
Published: (2025)
End-to-End On-Device Quantization-Aware Training for LLMs at Inference Cost
by: Tan, Qitao, et al.
Published: (2025)
by: Tan, Qitao, et al.
Published: (2025)
DAQ: Density-Aware Post-Training Weight-Only Quantization For LLMs
by: Luo, Yingsong, et al.
Published: (2024)
by: Luo, Yingsong, et al.
Published: (2024)
StableQAT: Stable Quantization-Aware Training at Ultra-Low Bitwidths
by: Chen, Tianyi, et al.
Published: (2026)
by: Chen, Tianyi, et al.
Published: (2026)
VQ-SAD: Vector Quantized Structure Aware Diffusion For Molecule Generation
by: Noravesh, Farshad, et al.
Published: (2026)
by: Noravesh, Farshad, et al.
Published: (2026)
DAQ: Delta-Aware Quantization for Post-Training LLM Weight Compression
by: Yu, Xiaoming, et al.
Published: (2026)
by: Yu, Xiaoming, et al.
Published: (2026)
RAMP: Reinforcement Adaptive Mixed Precision Quantization for Efficient On Device LLM Inference
by: Gautam, Arpit Singh, et al.
Published: (2026)
by: Gautam, Arpit Singh, et al.
Published: (2026)
Interpreting the Effects of Quantization on LLMs
by: Singh, Manpreet, et al.
Published: (2025)
by: Singh, Manpreet, et al.
Published: (2025)
LoTA-QAF: Lossless Ternary Adaptation for Quantization-Aware Fine-Tuning
by: Chen, Junyu, et al.
Published: (2025)
by: Chen, Junyu, et al.
Published: (2025)
Exact Learning of Arithmetic with Differentiable Agents
by: Papazov, Hristo, et al.
Published: (2025)
by: Papazov, Hristo, et al.
Published: (2025)
PhyPlan: Generalizable and Rapid Physical Task Planning with Physics Informed Skill Networks for Robot Manipulators
by: Chopra, Mudit, et al.
Published: (2024)
by: Chopra, Mudit, et al.
Published: (2024)
APAR: Modeling Irregular Target Functions in Tabular Regression via Arithmetic-Aware Pre-Training and Adaptive-Regularized Fine-Tuning
by: Wu, Hong-Wei, et al.
Published: (2024)
by: Wu, Hong-Wei, et al.
Published: (2024)
Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs
by: Kossen, Jannik, et al.
Published: (2024)
by: Kossen, Jannik, et al.
Published: (2024)
Efficient Edge LLMs Deployment via HessianAware Quantization and CPU GPU Collaborative
by: Zhang, Tuo, et al.
Published: (2025)
by: Zhang, Tuo, et al.
Published: (2025)
Scale When Needed: Adaptive Neuron-level Mixed Precision Quantization Aware Training
by: Varshney, Ayush K., et al.
Published: (2026)
by: Varshney, Ayush K., et al.
Published: (2026)
Low-Rank Quantization-Aware Training for LLMs
by: Bondarenko, Yelysei, et al.
Published: (2024)
by: Bondarenko, Yelysei, et al.
Published: (2024)
Error-Driven Prompt Optimization for Arithmetic Reasoning
by: Pándy, Árpád, et al.
Published: (2025)
by: Pándy, Árpád, et al.
Published: (2025)
RoSTE: An Efficient Quantization-Aware Supervised Fine-Tuning Approach for Large Language Models
by: Wei, Quan, et al.
Published: (2025)
by: Wei, Quan, et al.
Published: (2025)
Similar Items
-
MultiVer: Zero-Shot Multi-Agent Vulnerability Detection
by: Rajan, Shreshth
Published: (2026) -
TAGMol: Target-Aware Gradient-guided Molecule Generation
by: Dorna, Vineeth, et al.
Published: (2024) -
PeerRank: Autonomous LLM Evaluation Through Web-Grounded, Bias-Controlled Peer Review
by: Margalit, Yanki, et al.
Published: (2026) -
DQT: Dynamic Quantization Training via Dequantization-Free Nested Integer Arithmetic
by: Shalby, Hazem Hesham Yousef, et al.
Published: (2025) -
A Selective Quantization Tuner for ONNX Models
by: Louloudakis, Nikolaos, et al.
Published: (2025)