MobileQuant: Mobile-friendly Quantization for On-device Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Tan, Fuwen, Lee, Royson, Dudziak, Łukasz, Hu, Shell Xu, Bhattacharya, Sourav, Hospedales, Timothy, Tzimiropoulos, Georgios, Martinez, Brais |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A Bayesian Approach to Data Point Selection
di: Xu, Xinnuo, et al.
Pubblicazione: (2024)
di: Xu, Xinnuo, et al.
Pubblicazione: (2024)
Recurrent Early Exits for Federated Learning with Heterogeneous Clients
di: Lee, Royson, et al.
Pubblicazione: (2024)
di: Lee, Royson, et al.
Pubblicazione: (2024)
FedP$^2$EFT: Federated Learning to Personalize PEFT for Multilingual LLMs
di: Lee, Royson, et al.
Pubblicazione: (2025)
di: Lee, Royson, et al.
Pubblicazione: (2025)
Progressive Mixed-Precision Decoding for Efficient LLM Inference
di: Chen, Hao Mark, et al.
Pubblicazione: (2024)
di: Chen, Hao Mark, et al.
Pubblicazione: (2024)
CLIP-DPO: Vision-Language Models as a Source of Preference for Fixing Hallucinations in LVLMs
di: Ouali, Yassine, et al.
Pubblicazione: (2024)
di: Ouali, Yassine, et al.
Pubblicazione: (2024)
Feed-Forward Latent Domain Adaptation
di: Bohdal, Ondrej, et al.
Pubblicazione: (2022)
di: Bohdal, Ondrej, et al.
Pubblicazione: (2022)
Edge-SD-SR: Low Latency and Parameter Efficient On-device Super-Resolution with Stable Diffusion via Bidirectional Conditioning
di: Noroozi, Mehdi, et al.
Pubblicazione: (2024)
di: Noroozi, Mehdi, et al.
Pubblicazione: (2024)
OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models
di: Shao, Wenqi, et al.
Pubblicazione: (2023)
di: Shao, Wenqi, et al.
Pubblicazione: (2023)
Dynamic Expert Sharing: Decoupling Memory from Parallelism in Mixture-of-Experts Diffusion LLMs
di: Chen, Hao Mark, et al.
Pubblicazione: (2026)
di: Chen, Hao Mark, et al.
Pubblicazione: (2026)
Model Diffusion for Certifiable Few-shot Transfer Learning
di: Rezk, Fady, et al.
Pubblicazione: (2025)
di: Rezk, Fady, et al.
Pubblicazione: (2025)
Quant.npu: Enabling Efficient Mobile NPU Inference for on-device LLMs via Fully Static Quantization
di: Zhang, Jinghe, et al.
Pubblicazione: (2026)
di: Zhang, Jinghe, et al.
Pubblicazione: (2026)
FlatQuant: Flatness Matters for LLM Quantization
di: Sun, Yuxuan, et al.
Pubblicazione: (2024)
di: Sun, Yuxuan, et al.
Pubblicazione: (2024)
Graph Guided Question Answer Generation for Procedural Question-Answering
di: Pham, Hai X., et al.
Pubblicazione: (2024)
di: Pham, Hai X., et al.
Pubblicazione: (2024)
MambaQuant: Quantizing the Mamba Family with Variance Aligned Rotation Methods
di: Xu, Zukang, et al.
Pubblicazione: (2025)
di: Xu, Zukang, et al.
Pubblicazione: (2025)
You Only Need One Step: Fast Super-Resolution with Stable Diffusion via Scale Distillation
di: Noroozi, Mehdi, et al.
Pubblicazione: (2024)
di: Noroozi, Mehdi, et al.
Pubblicazione: (2024)
Knowledge Distillation Meets Open-Set Semi-Supervised Learning
di: Yang, Jing, et al.
Pubblicazione: (2022)
di: Yang, Jing, et al.
Pubblicazione: (2022)
Hierarchical Image Tokenization for Multi-Scale Image Super Resolution
di: Hadji, Isma, et al.
Pubblicazione: (2026)
di: Hadji, Isma, et al.
Pubblicazione: (2026)
PrefixQuant: Eliminating Outliers by Prefixed Tokens for Large Language Models Quantization
di: Chen, Mengzhao, et al.
Pubblicazione: (2024)
di: Chen, Mengzhao, et al.
Pubblicazione: (2024)
QuantMoE-Bench: Examining Post-Training Quantization for Mixture-of-Experts
di: Li, Pingzhi, et al.
Pubblicazione: (2024)
di: Li, Pingzhi, et al.
Pubblicazione: (2024)
Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs
di: Lu, Haiquan, et al.
Pubblicazione: (2026)
di: Lu, Haiquan, et al.
Pubblicazione: (2026)
FrameQuant: Flexible Low-Bit Quantization for Transformers
di: Adepu, Harshavardhan, et al.
Pubblicazione: (2024)
di: Adepu, Harshavardhan, et al.
Pubblicazione: (2024)
pQuant: Towards Effective Low-Bit Language Models via Decoupled Linear Quantization-Aware Training
di: Zhang, Wenzheng, et al.
Pubblicazione: (2026)
di: Zhang, Wenzheng, et al.
Pubblicazione: (2026)
SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
di: Xiao, Guangxuan, et al.
Pubblicazione: (2022)
di: Xiao, Guangxuan, et al.
Pubblicazione: (2022)
ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference
di: Liang, Yesheng, et al.
Pubblicazione: (2025)
di: Liang, Yesheng, et al.
Pubblicazione: (2025)
DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMs
di: Lin, Haokun, et al.
Pubblicazione: (2024)
di: Lin, Haokun, et al.
Pubblicazione: (2024)
DartQuant: Efficient Rotational Distribution Calibration for LLM Quantization
di: Shao, Yuantian, et al.
Pubblicazione: (2025)
di: Shao, Yuantian, et al.
Pubblicazione: (2025)
Aligned Unsupervised Pretraining of Object Detectors with Self-training
di: Metaxas, Ioannis Maniadis, et al.
Pubblicazione: (2023)
di: Metaxas, Ioannis Maniadis, et al.
Pubblicazione: (2023)
Fast Sampling Through The Reuse Of Attention Maps In Diffusion Models
di: Hunter, Rosco, et al.
Pubblicazione: (2023)
di: Hunter, Rosco, et al.
Pubblicazione: (2023)
FW-Merging: Scaling Model Merging with Frank-Wolfe Optimization
di: Chen, Hao Mark, et al.
Pubblicazione: (2025)
di: Chen, Hao Mark, et al.
Pubblicazione: (2025)
Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning
di: Zhao, Bingchen, et al.
Pubblicazione: (2024)
di: Zhao, Bingchen, et al.
Pubblicazione: (2024)
LIMP: Large Language Model Enhanced Intent-aware Mobility Prediction
di: Li, Songwei, et al.
Pubblicazione: (2024)
di: Li, Songwei, et al.
Pubblicazione: (2024)
Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents
di: Deng, Shihan, et al.
Pubblicazione: (2024)
di: Deng, Shihan, et al.
Pubblicazione: (2024)
Self-Supervised Multimodal Learning: A Survey
di: Zong, Yongshuo, et al.
Pubblicazione: (2023)
di: Zong, Yongshuo, et al.
Pubblicazione: (2023)
Mobile-MMLU: A Mobile Intelligence Language Understanding Benchmark
di: Bsharat, Sondos Mahmoud, et al.
Pubblicazione: (2025)
di: Bsharat, Sondos Mahmoud, et al.
Pubblicazione: (2025)
MUStReason: A Benchmark for Diagnosing Pragmatic Reasoning in Video-LMs for Multimodal Sarcasm Detection
di: Saha, Anisha, et al.
Pubblicazione: (2025)
di: Saha, Anisha, et al.
Pubblicazione: (2025)
ButterflyQuant: Ultra-low-bit LLM Quantization through Learnable Orthogonal Butterfly Transforms
di: Xu, Bingxin, et al.
Pubblicazione: (2025)
di: Xu, Bingxin, et al.
Pubblicazione: (2025)
GS-Quant: Granular Semantic and Generative Structural Quantization for Knowledge Graph Completion
di: Xie, Qizhuo, et al.
Pubblicazione: (2026)
di: Xie, Qizhuo, et al.
Pubblicazione: (2026)
Multi-scale Image Super Resolution with a Single Auto-Regressive Model
di: Sanchez, Enrique, et al.
Pubblicazione: (2025)
di: Sanchez, Enrique, et al.
Pubblicazione: (2025)
Agile-Quant: Activation-Guided Quantization for Faster Inference of LLMs on the Edge
di: Shen, Xuan, et al.
Pubblicazione: (2023)
di: Shen, Xuan, et al.
Pubblicazione: (2023)
Breaking Physical and Linguistic Borders: Multilingual Federated Prompt Tuning for Low-Resource Languages
di: Zhao, Wanru, et al.
Pubblicazione: (2025)
di: Zhao, Wanru, et al.
Pubblicazione: (2025)
Documenti analoghi
-
A Bayesian Approach to Data Point Selection
di: Xu, Xinnuo, et al.
Pubblicazione: (2024) -
Recurrent Early Exits for Federated Learning with Heterogeneous Clients
di: Lee, Royson, et al.
Pubblicazione: (2024) -
FedP$^2$EFT: Federated Learning to Personalize PEFT for Multilingual LLMs
di: Lee, Royson, et al.
Pubblicazione: (2025) -
Progressive Mixed-Precision Decoding for Efficient LLM Inference
di: Chen, Hao Mark, et al.
Pubblicazione: (2024) -
CLIP-DPO: Vision-Language Models as a Source of Preference for Fixing Hallucinations in LVLMs
di: Ouali, Yassine, et al.
Pubblicazione: (2024)