MobileLLM-Flash: Latency-Guided On-Device LLM Design for Industry Scale Deployment
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Hanxian, Fedorov, Igor, Gromov, Andrey, Beckerman, Bernard, Suda, Naveen, Eriksson, David, Balandat, Maximilian, Conway, Rylan, Huber, Patrick, Sankar, Chinnadhurai, Dalmia, Ayushi, Liu, Zechun, Wu, Lemeng, Elgamal, Tarek, Sagar, Adithya, Chandra, Vikas, Krishnamoorthi, Raghuraman |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MobileLLM-Pro Technical Report
by: Huber, Patrick, et al.
Published: (2025)
by: Huber, Patrick, et al.
Published: (2025)
MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases
by: Liu, Zechun, et al.
Published: (2024)
by: Liu, Zechun, et al.
Published: (2024)
MobileLLM-R1: Exploring the Limits of Sub-Billion Language Model Reasoners with Open Training Recipes
by: Zhao, Changsheng, et al.
Published: (2025)
by: Zhao, Changsheng, et al.
Published: (2025)
Short Data, Long Context: Distilling Positional Knowledge in Transformers
by: Huber, Patrick, et al.
Published: (2026)
by: Huber, Patrick, et al.
Published: (2026)
SpinQuant: LLM quantization with learned rotations
by: Liu, Zechun, et al.
Published: (2024)
by: Liu, Zechun, et al.
Published: (2024)
MobileMoE: Scaling On-Device Mixture of Experts
by: Chen, Yanbei, et al.
Published: (2026)
by: Chen, Yanbei, et al.
Published: (2026)
Llama Guard 3-1B-INT4: Compact and Efficient Safeguard for Human-AI Conversations
by: Fedorov, Igor, et al.
Published: (2024)
by: Fedorov, Igor, et al.
Published: (2024)
dTRPO: Trajectory Reduction in Policy Optimization of Diffusion Large Language Models
by: Zhang, Wenxuan, et al.
Published: (2026)
by: Zhang, Wenxuan, et al.
Published: (2026)
PathFusion: Path-consistent Lidar-Camera Deep Feature Fusion
by: Wu, Lemeng, et al.
Published: (2022)
by: Wu, Lemeng, et al.
Published: (2022)
ParetoQ: Improving Scaling Laws in Extremely Low-bit LLM Quantization
by: Liu, Zechun, et al.
Published: (2025)
by: Liu, Zechun, et al.
Published: (2025)
SqueezeSAM: User friendly mobile interactive segmentation
by: Varadarajan, Balakrishnan, et al.
Published: (2023)
by: Varadarajan, Balakrishnan, et al.
Published: (2023)
Efficient Track Anything
by: Xiong, Yunyang, et al.
Published: (2024)
by: Xiong, Yunyang, et al.
Published: (2024)
CoSMoEs: Compact Sparse Mixture of Experts
by: Huber, Patrick, et al.
Published: (2025)
by: Huber, Patrick, et al.
Published: (2025)
Communication Efficient Distributed Training with Distributed Lion
by: Liu, Bo, et al.
Published: (2024)
by: Liu, Bo, et al.
Published: (2024)
Small Vision-Language Models are Smart Compressors for Long Video Understanding
by: Fei, Junjie, et al.
Published: (2026)
by: Fei, Junjie, et al.
Published: (2026)
EdgeTAM: On-Device Track Anything Model
by: Zhou, Chong, et al.
Published: (2025)
by: Zhou, Chong, et al.
Published: (2025)
Informed Initialization for Bayesian Optimization and Active Learning
by: Hvarfner, Carl, et al.
Published: (2025)
by: Hvarfner, Carl, et al.
Published: (2025)
BONSAI: Bayesian Optimization with Natural Simplicity and Interpretability
by: Daulton, Samuel, et al.
Published: (2026)
by: Daulton, Samuel, et al.
Published: (2026)
LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding
by: Shen, Xiaoqian, et al.
Published: (2024)
by: Shen, Xiaoqian, et al.
Published: (2024)
WinQ: Accelerating Quantization-Aware Training of Language Models Around Saddle Points
by: Li, Dongyue, et al.
Published: (2026)
by: Li, Dongyue, et al.
Published: (2026)
Agent-as-a-Judge: Evaluate Agents with Agents
by: Zhuge, Mingchen, et al.
Published: (2024)
by: Zhuge, Mingchen, et al.
Published: (2024)
A poverty of reason : Sustainable development and economic growth / Wilfred Beckerman
by: Beckerman, Wilfred
by: Beckerman, Wilfred
How small should an economy's fiscal deficit be? : a monetary programming approach / Paul Beckerman
by: Beckerman, Paul
Published: (2000)
by: Beckerman, Paul
Published: (2000)
Optimal Foraging Group Size for a Human Population : The Case of Bari Fishing. / Stephen Beckerman
by: Beckerman, Stephen
Published: (1983)
by: Beckerman, Stephen
Published: (1983)
Towards Realistic Guarantees: A Probabilistic Certificate for SmoothLLM
by: Kumarappan, Adarsh, et al.
Published: (2025)
by: Kumarappan, Adarsh, et al.
Published: (2025)
Unexpected Improvements to Expected Improvement for Bayesian Optimization
by: Ament, Sebastian, et al.
Published: (2023)
by: Ament, Sebastian, et al.
Published: (2023)
PrE-Text: Training Language Models on Private Federated Data in the Age of LLMs
by: Hou, Charlie, et al.
Published: (2024)
by: Hou, Charlie, et al.
Published: (2024)
Mixture-of-Supernets: Improving Weight-Sharing Supernet Training with Architecture-Routed Mixture-of-Experts
by: Jawahar, Ganesh, et al.
Published: (2023)
by: Jawahar, Ganesh, et al.
Published: (2023)
DepthShrinker: A New Compression Paradigm Towards Boosting Real-Hardware Efficiency of Compact Neural Networks
by: Fu, Yonggan, et al.
Published: (2022)
by: Fu, Yonggan, et al.
Published: (2022)
Corded Ware Coastal Communities
by: Mariët Beckerman, Sandra
Published: (2021)
by: Mariët Beckerman, Sandra
Published: (2021)
Impact of the Open University on Public Libraries
by: Beckerman, Edwin P.
Published: (1975)
by: Beckerman, Edwin P.
Published: (1975)
The Woodbridge Tutorial Program
by: Beckerman, Edwin P.
Published: (1976)
by: Beckerman, Edwin P.
Published: (1976)
CoDi: Conversational Distillation for Grounded Question Answering
by: Huber, Patrick, et al.
Published: (2024)
by: Huber, Patrick, et al.
Published: (2024)
Towards LLM-Powered Verilog RTL Assistant: Self-Verification and Self-Correction
by: Huang, Hanxian, et al.
Published: (2024)
by: Huang, Hanxian, et al.
Published: (2024)
Robust Gaussian Processes via Relevance Pursuit
by: Ament, Sebastian, et al.
Published: (2024)
by: Ament, Sebastian, et al.
Published: (2024)
Dynamical Solution to the Eta Problem in Spectator Field Models
by: Elgamal, Sana, et al.
Published: (2025)
by: Elgamal, Sana, et al.
Published: (2025)
The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?
by: Banerjee, Sourav, et al.
Published: (2024)
by: Banerjee, Sourav, et al.
Published: (2024)
OrcaLoca: An LLM Agent Framework for Software Issue Localization
by: Yu, Zhongming, et al.
Published: (2025)
by: Yu, Zhongming, et al.
Published: (2025)
Script Gap: Evaluating LLM Triage on Indian Languages in Native vs Romanized Scripts in a Real World Setting
by: Khullar, Manurag, et al.
Published: (2025)
by: Khullar, Manurag, et al.
Published: (2025)
Is Our Chatbot Telling Lies? Assessing Correctness of an LLM-based Dutch Support Chatbot
by: Lassche, Herman, et al.
Published: (2024)
by: Lassche, Herman, et al.
Published: (2024)
Similar Items
-
MobileLLM-Pro Technical Report
by: Huber, Patrick, et al.
Published: (2025) -
MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases
by: Liu, Zechun, et al.
Published: (2024) -
MobileLLM-R1: Exploring the Limits of Sub-Billion Language Model Reasoners with Open Training Recipes
by: Zhao, Changsheng, et al.
Published: (2025) -
Short Data, Long Context: Distilling Positional Knowledge in Transformers
by: Huber, Patrick, et al.
Published: (2026) -
SpinQuant: LLM quantization with learned rotations
by: Liu, Zechun, et al.
Published: (2024)