A Systematic Evaluation of On-Device LLMs: Quantization, Performance, and Resources
Fuente:
arXiv
Guardado en:
| Autores principales: | Song, Qingyu, Liu, Rui, Lin, Wei, Liao, Peiyu, Zhao, Wenqian, Wang, Yiwen, Hu, Shoubo, Jiang, Yining, Long, Mochun, Zhen, Hui-Ling, Jiang, Ning, Yuan, Mingxuan, Xiang, Qiao, Xu, Hong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Innovation Discovery System for Networking Research
por: Zhang, Mengrui, et al.
Publicado: (2026)
por: Zhang, Mengrui, et al.
Publicado: (2026)
ReThinker: Scientific Reasoning by Rethinking with Guided Reflection and Confidence Control
por: Tang, Zhentao, et al.
Publicado: (2026)
por: Tang, Zhentao, et al.
Publicado: (2026)
One-Token Rollout: Guiding Supervised Fine-Tuning of LLMs with Policy Gradient
por: Ming, Rui, et al.
Publicado: (2025)
por: Ming, Rui, et al.
Publicado: (2025)
Inducing Systematicity in Transformers by Attending to Structurally Quantized Embeddings
por: Jiang, Yichen, et al.
Publicado: (2024)
por: Jiang, Yichen, et al.
Publicado: (2024)
ERVQ: Enhanced Residual Vector Quantization with Intra-and-Inter-Codebook Optimization for Neural Audio Codecs
por: Zheng, Rui-Chen, et al.
Publicado: (2024)
por: Zheng, Rui-Chen, et al.
Publicado: (2024)
MixPE: Quantization and Hardware Co-design for Efficient LLM Inference
por: Zhang, Yu, et al.
Publicado: (2024)
por: Zhang, Yu, et al.
Publicado: (2024)
The Regulation of Local Li + Coordination Environment for High‐Performance Quasi‐Solid‐State Polymer Electrolyte
por: Fan Yang, et al.
Publicado: (2025)
por: Fan Yang, et al.
Publicado: (2025)
A Streamable Neural Audio Codec with Residual Scalar-Vector Quantization for Real-Time Communication
por: Jiang, Xiao-Hang, et al.
Publicado: (2025)
por: Jiang, Xiao-Hang, et al.
Publicado: (2025)
QFT: Quantized Full-parameter Tuning of LLMs with Affordable Resources
por: Li, Zhikai, et al.
Publicado: (2023)
por: Li, Zhikai, et al.
Publicado: (2023)
MOSS: Efficient and Accurate FP8 LLM Training with Microscaling and Automatic Scaling
por: Zhang, Yu, et al.
Publicado: (2025)
por: Zhang, Yu, et al.
Publicado: (2025)
Greening the Grid: Electricity Market Clearing with Consumer-Based Carbon Cost
por: Jiang, Wenqian, et al.
Publicado: (2025)
por: Jiang, Wenqian, et al.
Publicado: (2025)
Intelligent Icing Detection Model of Wind Turbine Blades Based on SCADA data
por: Jiang, Wenqian, et al.
Publicado: (2021)
por: Jiang, Wenqian, et al.
Publicado: (2021)
End-to-End On-Device Quantization-Aware Training for LLMs at Inference Cost
por: Tan, Qitao, et al.
Publicado: (2025)
por: Tan, Qitao, et al.
Publicado: (2025)
SVDq: 1.25-bit and 410x Key Cache Compression for LLM Attention
por: Yankun, Hong, et al.
Publicado: (2025)
por: Yankun, Hong, et al.
Publicado: (2025)
Resource-Efficient Teleportation of High-Dimensional Quantum Coherence via Initial Phase Engineering
por: Huang, Long, et al.
Publicado: (2026)
por: Huang, Long, et al.
Publicado: (2026)
Nanoconjugate Improves Cognitive Deficit and Limits the Pathogenic Tau Burden in Okadaic‐Acid‐Induced Alzheimer's Mice
por: Qiuju Liang, et al.
Publicado: (2025)
por: Qiuju Liang, et al.
Publicado: (2025)
Quantization Meets dLLMs: A Systematic Study of Post-training Quantization for Diffusion LLMs
por: Lin, Haokun, et al.
Publicado: (2025)
por: Lin, Haokun, et al.
Publicado: (2025)
Empirical Guidelines for Deploying LLMs onto Resource-constrained Edge Devices
por: Qin, Ruiyang, et al.
Publicado: (2024)
por: Qin, Ruiyang, et al.
Publicado: (2024)
PQD: Post-training Quantization for Efficient Diffusion Models
por: Ye, Jiaojiao, et al.
Publicado: (2024)
por: Ye, Jiaojiao, et al.
Publicado: (2024)
Embers of Active Galactic Nuclei: Tidal Disruption Events and Quasiperiodic Eruptions
por: Jiang, Ning, et al.
Publicado: (2025)
por: Jiang, Ning, et al.
Publicado: (2025)
KVTuner: Sensitivity-Aware Layer-Wise Mixed-Precision KV Cache Quantization for Efficient and Nearly Lossless LLM Inference
por: Li, Xing, et al.
Publicado: (2025)
por: Li, Xing, et al.
Publicado: (2025)
Cell‐Specific Control of Mammalian Gene Expression Using DNA Repair Inducible Ribozyme Switches
por: Jieling Hong, et al.
Publicado: (2024)
por: Jieling Hong, et al.
Publicado: (2024)
Cell‐Specific Control of Mammalian Gene Expression Using DNA Repair Inducible Ribozyme Switches
por: Jieling Hong, et al.
Publicado: (2024)
por: Jieling Hong, et al.
Publicado: (2024)
Sharp upper bounds on the $A_α$-spectral radius of graphs
por: Hong, Zhen-Mu, et al.
Publicado: (2026)
por: Hong, Zhen-Mu, et al.
Publicado: (2026)
On the disjunctive domination numbers of the torus grid graphs
por: Qiao, Zhi, et al.
Publicado: (2026)
por: Qiao, Zhi, et al.
Publicado: (2026)
BetterV: Controlled Verilog Generation with Discriminative Guidance
por: Pei, Zehua, et al.
Publicado: (2024)
por: Pei, Zehua, et al.
Publicado: (2024)
Awakening of A Blazar at Redshift 2.7 Temporally Coincident with Arrival of Cospatial Neutrino Event IceCube-201221A
por: Jiang, Xiong, et al.
Publicado: (2024)
por: Jiang, Xiong, et al.
Publicado: (2024)
HBLLM: Wavelet-Enhanced High-Fidelity 1-Bit Quantization for LLMs
por: Chen, Ningning, et al.
Publicado: (2025)
por: Chen, Ningning, et al.
Publicado: (2025)
Systematic study of capture thresholds with time dependent Hartree-Fock theory
por: Yao, Hong, et al.
Publicado: (2024)
por: Yao, Hong, et al.
Publicado: (2024)
Hypercrosslinked porous and coordination polymer materials for electrolyte membranes in lithium‐metal batteries
por: Mochun Zhang, et al.
Publicado: (2024)
por: Mochun Zhang, et al.
Publicado: (2024)
Certifying Language Model Robustness with Fuzzed Randomized Smoothing: An Efficient Defense Against Backdoor Attacks
por: He, Bowei, et al.
Publicado: (2025)
por: He, Bowei, et al.
Publicado: (2025)
Uncovering Cross-Objective Interference in Multi-Objective Alignment
por: Lu, Yining, et al.
Publicado: (2026)
por: Lu, Yining, et al.
Publicado: (2026)
MD-AirComp+: Adaptive Quantization for Blind Massive Digital Over-the-Air Computation
por: Qiao, Li, et al.
Publicado: (2026)
por: Qiao, Li, et al.
Publicado: (2026)
QKVShare: Quantized KV-Cache Handoff for Multi-Agent On-Device LLMs
por: Honavar, Pratik, et al.
Publicado: (2026)
por: Honavar, Pratik, et al.
Publicado: (2026)
Beyond Speedup -- Utilizing KV Cache for Sampling and Reasoning
por: Xing, Zeyu, et al.
Publicado: (2026)
por: Xing, Zeyu, et al.
Publicado: (2026)
Consumer-based Carbon Costs: Integrating Consumer Carbon Preferences in Electricity Markets
por: Jiang, Wenqian, et al.
Publicado: (2025)
por: Jiang, Wenqian, et al.
Publicado: (2025)
DAQ: Density-Aware Post-Training Weight-Only Quantization For LLMs
por: Luo, Yingsong, et al.
Publicado: (2024)
por: Luo, Yingsong, et al.
Publicado: (2024)
DemoTuner: Automatic Performance Tuning for Database Management Systems Based on Demonstration Reinforcement Learning
por: Dou, Hui, et al.
Publicado: (2025)
por: Dou, Hui, et al.
Publicado: (2025)
Quantization-Aware Neuromorphic Architecture for Skin Disease Classification on Resource-Constrained Devices
por: Wang, Haitian, et al.
Publicado: (2025)
por: Wang, Haitian, et al.
Publicado: (2025)
PSD: Pushing the Pareto Frontier of Diffusion LLMs via Parallel Speculative Decoding
por: Sun, Shengyin, et al.
Publicado: (2026)
por: Sun, Shengyin, et al.
Publicado: (2026)
Ejemplares similares
-
Innovation Discovery System for Networking Research
por: Zhang, Mengrui, et al.
Publicado: (2026) -
ReThinker: Scientific Reasoning by Rethinking with Guided Reflection and Confidence Control
por: Tang, Zhentao, et al.
Publicado: (2026) -
One-Token Rollout: Guiding Supervised Fine-Tuning of LLMs with Policy Gradient
por: Ming, Rui, et al.
Publicado: (2025) -
Inducing Systematicity in Transformers by Attending to Structurally Quantized Embeddings
por: Jiang, Yichen, et al.
Publicado: (2024) -
ERVQ: Enhanced Residual Vector Quantization with Intra-and-Inter-Codebook Optimization for Neural Audio Codecs
por: Zheng, Rui-Chen, et al.
Publicado: (2024)