Saved in:
| Main Authors: | Yadav, Agatsya, Bhargavi, Renta Chintala |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2512.06490 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs
by: Noh, Kanghyun, et al.
Published: (2026)
by: Noh, Kanghyun, et al.
Published: (2026)
CPTQuant - A Novel Mixed Precision Post-Training Quantization Techniques for Large Language Models
by: Nanda, Amitash, et al.
Published: (2024)
by: Nanda, Amitash, et al.
Published: (2024)
FedMM-X: A Trustworthy and Interpretable Framework for Federated Multi-Modal Learning in Dynamic Environments
by: Balija, Sree Bhargavi
Published: (2025)
by: Balija, Sree Bhargavi
Published: (2025)
DynaMem: Online Dynamic Spatio-Semantic Memory for Open World Mobile Manipulation
by: Liu, Peiqi, et al.
Published: (2024)
by: Liu, Peiqi, et al.
Published: (2024)
RSQ: Learning from Important Tokens Leads to Better Quantized LLMs
by: Sung, Yi-Lin, et al.
Published: (2025)
by: Sung, Yi-Lin, et al.
Published: (2025)
Two-Stage Grid Optimization for Group-wise Quantization of LLMs
by: Kim, Junhan, et al.
Published: (2026)
by: Kim, Junhan, et al.
Published: (2026)
ELUTQ: Optimizing Quantization Accuracy under LUT-Based Computation for Edge LLMs
by: Nie, Xin, et al.
Published: (2025)
by: Nie, Xin, et al.
Published: (2025)
Pyramid Vector Quantization for LLMs
by: van der Ouderaa, Tycho F. A., et al.
Published: (2024)
by: van der Ouderaa, Tycho F. A., et al.
Published: (2024)
Quant.npu: Enabling Efficient Mobile NPU Inference for on-device LLMs via Fully Static Quantization
by: Zhang, Jinghe, et al.
Published: (2026)
by: Zhang, Jinghe, et al.
Published: (2026)
ConQuR: Corner Aligned Activation Quantization via Optimized Rotations for LLMs
by: Thrash, Chayne, et al.
Published: (2026)
by: Thrash, Chayne, et al.
Published: (2026)
FedNAMs: Performing Interpretability Analysis in Federated Learning Context
by: Nanda, Amitash, et al.
Published: (2025)
by: Nanda, Amitash, et al.
Published: (2025)
Low-Rank Correction for Quantized LLMs
by: Scetbon, Meyer, et al.
Published: (2024)
by: Scetbon, Meyer, et al.
Published: (2024)
Stressor Type Matters! -- Exploring Factors Influencing Cross-Dataset Generalizability of Physiological Stress Detection
by: Prajod, Pooja, et al.
Published: (2024)
by: Prajod, Pooja, et al.
Published: (2024)
Trace is the Next AutoDiff: Generative Optimization with Rich Feedback, Execution Traces, and LLMs
by: Cheng, Ching-An, et al.
Published: (2024)
by: Cheng, Ching-An, et al.
Published: (2024)
Metis: Training LLMs with FP4 Quantization
by: Cao, Hengjie, et al.
Published: (2025)
by: Cao, Hengjie, et al.
Published: (2025)
Optimize Weight Rounding via Signed Gradient Descent for the Quantization of LLMs
by: Cheng, Wenhua, et al.
Published: (2023)
by: Cheng, Wenhua, et al.
Published: (2023)
YOR: Your Own Mobile Manipulator for Generalizable Robotics
by: Anjaria, Manan H, et al.
Published: (2026)
by: Anjaria, Manan H, et al.
Published: (2026)
Layer-Wise Quantization: A Pragmatic and Effective Method for Quantizing LLMs Beyond Integer Bit-Levels
by: Dumitru, Razvan-Gabriel, et al.
Published: (2024)
by: Dumitru, Razvan-Gabriel, et al.
Published: (2024)
Low-Bit Quantization Favors Undertrained LLMs: Scaling Laws for Quantized LLMs with 100T Training Tokens
by: Ouyang, Xu, et al.
Published: (2024)
by: Ouyang, Xu, et al.
Published: (2024)
Executable Counterfactuals: Improving LLMs' Causal Reasoning Through Code
by: Vashishtha, Aniket, et al.
Published: (2025)
by: Vashishtha, Aniket, et al.
Published: (2025)
Preserve-Then-Quantize: Balancing Rank Budgets for Quantization Error Reconstruction in LLMs
by: Cho, Yoonjun, et al.
Published: (2026)
by: Cho, Yoonjun, et al.
Published: (2026)
Autoformulation of Mathematical Optimization Models Using LLMs
by: Astorga, Nicolás, et al.
Published: (2024)
by: Astorga, Nicolás, et al.
Published: (2024)
Decoding Federated Learning: The FedNAM+ Conformal Revolution
by: Balija, Sree Bhargavi, et al.
Published: (2025)
by: Balija, Sree Bhargavi, et al.
Published: (2025)
Interpreting the Effects of Quantization on LLMs
by: Singh, Manpreet, et al.
Published: (2025)
by: Singh, Manpreet, et al.
Published: (2025)
Optimizing Large Language Model Training Using FP4 Quantization
by: Wang, Ruizhe, et al.
Published: (2025)
by: Wang, Ruizhe, et al.
Published: (2025)
Effective Quantization of Muon Optimizer States
by: Gupta, Aman, et al.
Published: (2025)
by: Gupta, Aman, et al.
Published: (2025)
Husky: A Unified, Open-Source Language Agent for Multi-Step Reasoning
by: Kim, Joongwon, et al.
Published: (2024)
by: Kim, Joongwon, et al.
Published: (2024)
QEFT: Quantization for Efficient Fine-Tuning of LLMs
by: Lee, Changhun, et al.
Published: (2024)
by: Lee, Changhun, et al.
Published: (2024)
How Does Quantization Affect Multilingual LLMs?
by: Marchisio, Kelly, et al.
Published: (2024)
by: Marchisio, Kelly, et al.
Published: (2024)
A Systematic Evaluation of On-Device LLMs: Quantization, Performance, and Resources
by: Song, Qingyu, et al.
Published: (2025)
by: Song, Qingyu, et al.
Published: (2025)
CCQ: Convolutional Code for Extreme Low-bit Quantization in LLMs
by: Zhou, Zhaojing, et al.
Published: (2025)
by: Zhou, Zhaojing, et al.
Published: (2025)
Understanding the Difficulty of Low-Precision Post-Training Quantization for LLMs
by: Xu, Zifei, et al.
Published: (2024)
by: Xu, Zifei, et al.
Published: (2024)
Traffic and Mobility Optimization Using AI: Comparative Study between Dubai and Riyadh
by: Aalijah, Kanwal
Published: (2025)
by: Aalijah, Kanwal
Published: (2025)
A Predictive and Optimization Approach for Enhanced Urban Mobility Using Spatiotemporal Data
by: Mishra, Shambhavi, et al.
Published: (2024)
by: Mishra, Shambhavi, et al.
Published: (2024)
Advancing Audio Fingerprinting Accuracy Addressing Background Noise and Distortion Challenges
by: Kamuni, Navin, et al.
Published: (2024)
by: Kamuni, Navin, et al.
Published: (2024)
Training-Inference Consistent Segmented Execution for Long-Context LLMs
by: Shang, Xianpeng, et al.
Published: (2026)
by: Shang, Xianpeng, et al.
Published: (2026)
Efficient Dynamic and Momentum Aperture Optimization for Lattice Design Using Multipoint Bayesian Algorithm Execution
by: Zhang, Z., et al.
Published: (2025)
by: Zhang, Z., et al.
Published: (2025)
On-the-Fly Adaptation to Quantization: Configuration-Aware LoRA for Efficient Fine-Tuning of Quantized LLMs
by: Ye, Rongguang, et al.
Published: (2025)
by: Ye, Rongguang, et al.
Published: (2025)
Quantized Evolution Strategies: High-precision Fine-tuning of Quantized LLMs at Low-precision Cost
by: Xu, Yinggan, et al.
Published: (2026)
by: Xu, Yinggan, et al.
Published: (2026)
EXAQ: Exponent Aware Quantization For LLMs Acceleration
by: Shkolnik, Moran, et al.
Published: (2024)
by: Shkolnik, Moran, et al.
Published: (2024)
Similar Items
-
QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs
by: Noh, Kanghyun, et al.
Published: (2026) -
CPTQuant - A Novel Mixed Precision Post-Training Quantization Techniques for Large Language Models
by: Nanda, Amitash, et al.
Published: (2024) -
FedMM-X: A Trustworthy and Interpretable Framework for Federated Multi-Modal Learning in Dynamic Environments
by: Balija, Sree Bhargavi
Published: (2025) -
DynaMem: Online Dynamic Spatio-Semantic Memory for Open World Mobile Manipulation
by: Liu, Peiqi, et al.
Published: (2024) -
RSQ: Learning from Important Tokens Leads to Better Quantized LLMs
by: Sung, Yi-Lin, et al.
Published: (2025)