Optimizing LLMs Using Quantization for Mobile Execution
Fuente:
arXiv
Guardado en:
| Autores principales: | Yadav, Agatsya, Bhargavi, Renta Chintala |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs
por: Noh, Kanghyun, et al.
Publicado: (2026)
por: Noh, Kanghyun, et al.
Publicado: (2026)
CPTQuant - A Novel Mixed Precision Post-Training Quantization Techniques for Large Language Models
por: Nanda, Amitash, et al.
Publicado: (2024)
por: Nanda, Amitash, et al.
Publicado: (2024)
FedMM-X: A Trustworthy and Interpretable Framework for Federated Multi-Modal Learning in Dynamic Environments
por: Balija, Sree Bhargavi
Publicado: (2025)
por: Balija, Sree Bhargavi
Publicado: (2025)
DynaMem: Online Dynamic Spatio-Semantic Memory for Open World Mobile Manipulation
por: Liu, Peiqi, et al.
Publicado: (2024)
por: Liu, Peiqi, et al.
Publicado: (2024)
ELUTQ: Optimizing Quantization Accuracy under LUT-Based Computation for Edge LLMs
por: Nie, Xin, et al.
Publicado: (2025)
por: Nie, Xin, et al.
Publicado: (2025)
Pyramid Vector Quantization for LLMs
por: van der Ouderaa, Tycho F. A., et al.
Publicado: (2024)
por: van der Ouderaa, Tycho F. A., et al.
Publicado: (2024)
Two-Stage Grid Optimization for Group-wise Quantization of LLMs
por: Kim, Junhan, et al.
Publicado: (2026)
por: Kim, Junhan, et al.
Publicado: (2026)
RSQ: Learning from Important Tokens Leads to Better Quantized LLMs
por: Sung, Yi-Lin, et al.
Publicado: (2025)
por: Sung, Yi-Lin, et al.
Publicado: (2025)
ConQuR: Corner Aligned Activation Quantization via Optimized Rotations for LLMs
por: Thrash, Chayne, et al.
Publicado: (2026)
por: Thrash, Chayne, et al.
Publicado: (2026)
Low-Rank Correction for Quantized LLMs
por: Scetbon, Meyer, et al.
Publicado: (2024)
por: Scetbon, Meyer, et al.
Publicado: (2024)
Quant.npu: Enabling Efficient Mobile NPU Inference for on-device LLMs via Fully Static Quantization
por: Zhang, Jinghe, et al.
Publicado: (2026)
por: Zhang, Jinghe, et al.
Publicado: (2026)
Trace is the Next AutoDiff: Generative Optimization with Rich Feedback, Execution Traces, and LLMs
por: Cheng, Ching-An, et al.
Publicado: (2024)
por: Cheng, Ching-An, et al.
Publicado: (2024)
Metis: Training LLMs with FP4 Quantization
por: Cao, Hengjie, et al.
Publicado: (2025)
por: Cao, Hengjie, et al.
Publicado: (2025)
Autoformulation of Mathematical Optimization Models Using LLMs
por: Astorga, Nicolás, et al.
Publicado: (2024)
por: Astorga, Nicolás, et al.
Publicado: (2024)
Executable Counterfactuals: Improving LLMs' Causal Reasoning Through Code
por: Vashishtha, Aniket, et al.
Publicado: (2025)
por: Vashishtha, Aniket, et al.
Publicado: (2025)
Low-Bit Quantization Favors Undertrained LLMs: Scaling Laws for Quantized LLMs with 100T Training Tokens
por: Ouyang, Xu, et al.
Publicado: (2024)
por: Ouyang, Xu, et al.
Publicado: (2024)
Optimize Weight Rounding via Signed Gradient Descent for the Quantization of LLMs
por: Cheng, Wenhua, et al.
Publicado: (2023)
por: Cheng, Wenhua, et al.
Publicado: (2023)
Preserve-Then-Quantize: Balancing Rank Budgets for Quantization Error Reconstruction in LLMs
por: Cho, Yoonjun, et al.
Publicado: (2026)
por: Cho, Yoonjun, et al.
Publicado: (2026)
Effective Quantization of Muon Optimizer States
por: Gupta, Aman, et al.
Publicado: (2025)
por: Gupta, Aman, et al.
Publicado: (2025)
FedNAMs: Performing Interpretability Analysis in Federated Learning Context
por: Nanda, Amitash, et al.
Publicado: (2025)
por: Nanda, Amitash, et al.
Publicado: (2025)
Optimizing Large Language Model Training Using FP4 Quantization
por: Wang, Ruizhe, et al.
Publicado: (2025)
por: Wang, Ruizhe, et al.
Publicado: (2025)
A Systematic Evaluation of On-Device LLMs: Quantization, Performance, and Resources
por: Song, Qingyu, et al.
Publicado: (2025)
por: Song, Qingyu, et al.
Publicado: (2025)
CCQ: Convolutional Code for Extreme Low-bit Quantization in LLMs
por: Zhou, Zhaojing, et al.
Publicado: (2025)
por: Zhou, Zhaojing, et al.
Publicado: (2025)
Traffic and Mobility Optimization Using AI: Comparative Study between Dubai and Riyadh
por: Aalijah, Kanwal
Publicado: (2025)
por: Aalijah, Kanwal
Publicado: (2025)
Understanding the Difficulty of Low-Precision Post-Training Quantization for LLMs
por: Xu, Zifei, et al.
Publicado: (2024)
por: Xu, Zifei, et al.
Publicado: (2024)
A Predictive and Optimization Approach for Enhanced Urban Mobility Using Spatiotemporal Data
por: Mishra, Shambhavi, et al.
Publicado: (2024)
por: Mishra, Shambhavi, et al.
Publicado: (2024)
Stressor Type Matters! -- Exploring Factors Influencing Cross-Dataset Generalizability of Physiological Stress Detection
por: Prajod, Pooja, et al.
Publicado: (2024)
por: Prajod, Pooja, et al.
Publicado: (2024)
Interpreting the Effects of Quantization on LLMs
por: Singh, Manpreet, et al.
Publicado: (2025)
por: Singh, Manpreet, et al.
Publicado: (2025)
YOR: Your Own Mobile Manipulator for Generalizable Robotics
por: Anjaria, Manan H, et al.
Publicado: (2026)
por: Anjaria, Manan H, et al.
Publicado: (2026)
QEFT: Quantization for Efficient Fine-Tuning of LLMs
por: Lee, Changhun, et al.
Publicado: (2024)
por: Lee, Changhun, et al.
Publicado: (2024)
How Does Quantization Affect Multilingual LLMs?
por: Marchisio, Kelly, et al.
Publicado: (2024)
por: Marchisio, Kelly, et al.
Publicado: (2024)
Layer-Wise Quantization: A Pragmatic and Effective Method for Quantizing LLMs Beyond Integer Bit-Levels
por: Dumitru, Razvan-Gabriel, et al.
Publicado: (2024)
por: Dumitru, Razvan-Gabriel, et al.
Publicado: (2024)
CLAQ: Pushing the Limits of Low-Bit Post-Training Quantization for LLMs
por: Wang, Haoyu, et al.
Publicado: (2024)
por: Wang, Haoyu, et al.
Publicado: (2024)
Efficient Dynamic and Momentum Aperture Optimization for Lattice Design Using Multipoint Bayesian Algorithm Execution
por: Zhang, Z., et al.
Publicado: (2025)
por: Zhang, Z., et al.
Publicado: (2025)
Training-Inference Consistent Segmented Execution for Long-Context LLMs
por: Shang, Xianpeng, et al.
Publicado: (2026)
por: Shang, Xianpeng, et al.
Publicado: (2026)
Leveraging RAG-LLMs for Urban Mobility Simulation and Analysis
por: Ding, Yue, et al.
Publicado: (2025)
por: Ding, Yue, et al.
Publicado: (2025)
On-the-Fly Adaptation to Quantization: Configuration-Aware LoRA for Efficient Fine-Tuning of Quantized LLMs
por: Ye, Rongguang, et al.
Publicado: (2025)
por: Ye, Rongguang, et al.
Publicado: (2025)
Quantized Evolution Strategies: High-precision Fine-tuning of Quantized LLMs at Low-precision Cost
por: Xu, Yinggan, et al.
Publicado: (2026)
por: Xu, Yinggan, et al.
Publicado: (2026)
Beyond Outliers: A Study of Optimizers Under Quantization
por: Vlassis, Georgios, et al.
Publicado: (2025)
por: Vlassis, Georgios, et al.
Publicado: (2025)
StatQAT: Statistical Quantizer Optimization for Deep Networks
por: Aktukmak, Mehmet, et al.
Publicado: (2026)
por: Aktukmak, Mehmet, et al.
Publicado: (2026)
Ejemplares similares
-
QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs
por: Noh, Kanghyun, et al.
Publicado: (2026) -
CPTQuant - A Novel Mixed Precision Post-Training Quantization Techniques for Large Language Models
por: Nanda, Amitash, et al.
Publicado: (2024) -
FedMM-X: A Trustworthy and Interpretable Framework for Federated Multi-Modal Learning in Dynamic Environments
por: Balija, Sree Bhargavi
Publicado: (2025) -
DynaMem: Online Dynamic Spatio-Semantic Memory for Open World Mobile Manipulation
por: Liu, Peiqi, et al.
Publicado: (2024) -
ELUTQ: Optimizing Quantization Accuracy under LUT-Based Computation for Edge LLMs
por: Nie, Xin, et al.
Publicado: (2025)