MCUBERT: Memory-Efficient BERT Inference on Commodity Microcontrollers
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Zebin, Chen, Renze, Wu, Taiqiang, Wong, Ngai, Liang, Yun, Wang, Runsheng, Huang, Ru, Li, Meng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mixture-of-Subspaces in Low-Rank Adaptation
by: Wu, Taiqiang, et al.
Published: (2024)
by: Wu, Taiqiang, et al.
Published: (2024)
LLM-NEO: Parameter Efficient Knowledge Distillation for Large Language Models
by: Yang, Runming, et al.
Published: (2024)
by: Yang, Runming, et al.
Published: (2024)
Edge-free but Structure-aware: Prototype-Guided Knowledge Distillation from GNNs to MLPs
by: Wu, Taiqiang, et al.
Published: (2023)
by: Wu, Taiqiang, et al.
Published: (2023)
TS-Memory: Plug-and-Play Memory for Time Series Foundation Models
by: Lyu, Sisuo, et al.
Published: (2026)
by: Lyu, Sisuo, et al.
Published: (2026)
FastQuery: Communication-efficient Embedding Table Query for Private LLM Inference
by: Lin, Chenqi, et al.
Published: (2024)
by: Lin, Chenqi, et al.
Published: (2024)
The Art of Efficient Reasoning: Data, Reward, and Optimization
by: Wu, Taiqiang, et al.
Published: (2026)
by: Wu, Taiqiang, et al.
Published: (2026)
AttentionLego: An Open-Source Building Block For Spatially-Scalable Large Language Model Accelerator With Processing-In-Memory Technology
by: Cong, Rongqing, et al.
Published: (2024)
by: Cong, Rongqing, et al.
Published: (2024)
Weight-Inherited Distillation for Task-Agnostic BERT Compression
by: Wu, Taiqiang, et al.
Published: (2023)
by: Wu, Taiqiang, et al.
Published: (2023)
Stochastic Multivariate Universal-Radix Finite-State Machine: a Theoretically and Practically Elegant Nonlinear Function Approximator
by: Feng, Xincheng, et al.
Published: (2024)
by: Feng, Xincheng, et al.
Published: (2024)
PDNNet: PDN-Aware GNN-CNN Heterogeneous Network for Dynamic IR Drop Prediction
by: Zhao, Yuxiang, et al.
Published: (2024)
by: Zhao, Yuxiang, et al.
Published: (2024)
Revisiting Model Interpolation for Efficient Reasoning
by: Wu, Taiqiang, et al.
Published: (2025)
by: Wu, Taiqiang, et al.
Published: (2025)
Evaluating the Energy Efficiency of NPU-Accelerated Machine Learning Inference on Embedded Microcontrollers
by: Fanariotis, Anastasios, et al.
Published: (2025)
by: Fanariotis, Anastasios, et al.
Published: (2025)
CSI-BERT2: A BERT-inspired Framework for Efficient CSI Prediction and Classification in Wireless Communication and Sensing
by: Zhao, Zijian, et al.
Published: (2024)
by: Zhao, Zijian, et al.
Published: (2024)
CERSA: Cumulative Energy-Retaining Subspace Adaptation for Memory-Efficient Fine-Tuning
by: Ge, Jingze, et al.
Published: (2026)
by: Ge, Jingze, et al.
Published: (2026)
HeadInfer: Memory-Efficient LLM Inference by Head-wise Offloading
by: Luo, Cheng, et al.
Published: (2025)
by: Luo, Cheng, et al.
Published: (2025)
Injecting linguistic knowledge into BERT for Dialogue State Tracking
by: Feng, Xiaohan, et al.
Published: (2023)
by: Feng, Xiaohan, et al.
Published: (2023)
LoRETTA: Low-Rank Economic Tensor-Train Adaptation for Ultra-Low-Parameter Fine-Tuning of Large Language Models
by: Yang, Yifan, et al.
Published: (2024)
by: Yang, Yifan, et al.
Published: (2024)
Regret-Guided Search Control for Efficient Learning in AlphaZero
by: Tsai, Yun-Jui, et al.
Published: (2026)
by: Tsai, Yun-Jui, et al.
Published: (2026)
GenoBERT: A Language Model for Accurate Genotype Imputation
by: Huang, Lei, et al.
Published: (2026)
by: Huang, Lei, et al.
Published: (2026)
vMCU: Coordinated Memory Management and Kernel Optimization for DNN Inference on MCUs
by: Zheng, Size, et al.
Published: (2024)
by: Zheng, Size, et al.
Published: (2024)
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models
by: Shao, Zishan, et al.
Published: (2025)
by: Shao, Zishan, et al.
Published: (2025)
BuddyMoE: Exploiting Expert Redundancy to Accelerate Memory-Constrained Mixture-of-Experts Inference
by: Wang, Yun, et al.
Published: (2025)
by: Wang, Yun, et al.
Published: (2025)
AIS: Adaptive Importance Sampling for Quantized RL
by: Zhou, Jiajun, et al.
Published: (2026)
by: Zhou, Jiajun, et al.
Published: (2026)
On-Device Training of Fully Quantized Deep Neural Networks on Cortex-M Microcontrollers
by: Deutel, Mark, et al.
Published: (2024)
by: Deutel, Mark, et al.
Published: (2024)
MELINOE: Fine-Tuning Enables Memory-Efficient Inference for Mixture-of-Experts Models
by: Raje, Arian, et al.
Published: (2026)
by: Raje, Arian, et al.
Published: (2026)
NeuZip: Memory-Efficient Training and Inference with Dynamic Compression of Neural Networks
by: Hao, Yongchang, et al.
Published: (2024)
by: Hao, Yongchang, et al.
Published: (2024)
On-the-Fly Adaptation to Quantization: Configuration-Aware LoRA for Efficient Fine-Tuning of Quantized LLMs
by: Ye, Rongguang, et al.
Published: (2025)
by: Ye, Rongguang, et al.
Published: (2025)
APTQ: Attention-aware Post-Training Mixed-Precision Quantization for Large Language Models
by: Guan, Ziyi, et al.
Published: (2024)
by: Guan, Ziyi, et al.
Published: (2024)
Memory- and Latency-Constrained Inference of Large Language Models via Adaptive Split Computing
by: Sung, Mingyu, et al.
Published: (2025)
by: Sung, Mingyu, et al.
Published: (2025)
AdapMoE: Adaptive Sensitivity-based Expert Gating and Management for Efficient MoE Inference
by: Zhong, Shuzhang, et al.
Published: (2024)
by: Zhong, Shuzhang, et al.
Published: (2024)
FedSI: Federated Subnetwork Inference for Efficient Uncertainty Quantification
by: Chen, Hui, et al.
Published: (2024)
by: Chen, Hui, et al.
Published: (2024)
Space-aware Socioeconomic Indicator Inference with Heterogeneous Graphs
by: Zou, Xingchen, et al.
Published: (2024)
by: Zou, Xingchen, et al.
Published: (2024)
Quant.npu: Enabling Efficient Mobile NPU Inference for on-device LLMs via Fully Static Quantization
by: Zhang, Jinghe, et al.
Published: (2026)
by: Zhang, Jinghe, et al.
Published: (2026)
KV Admission: Learning What to Write for Efficient Long-Context Inference
by: Huang, Yen-Chieh, et al.
Published: (2025)
by: Huang, Yen-Chieh, et al.
Published: (2025)
Synergizing Foundation Models and Federated Learning: A Survey
by: Li, Shenghui, et al.
Published: (2024)
by: Li, Shenghui, et al.
Published: (2024)
PTQTP: Post-Training Quantization to Trit-Planes for Large Language Models
by: Xiao, He, et al.
Published: (2025)
by: Xiao, He, et al.
Published: (2025)
4-bit Shampoo for Memory-Efficient Network Training
by: Wang, Sike, et al.
Published: (2024)
by: Wang, Sike, et al.
Published: (2024)
Optimal Bayesian Stopping for Efficient Inference of Consistent LLM Answers
by: Huang, Jingkai, et al.
Published: (2026)
by: Huang, Jingkai, et al.
Published: (2026)
Memory-Efficient LLM Training with Online Subspace Descent
by: Liang, Kaizhao, et al.
Published: (2024)
by: Liang, Kaizhao, et al.
Published: (2024)
FedKRSO: Communication and Memory Efficient Federated Fine-Tuning of Large Language Models
by: Yang, Guohao, et al.
Published: (2026)
by: Yang, Guohao, et al.
Published: (2026)
Similar Items
-
Mixture-of-Subspaces in Low-Rank Adaptation
by: Wu, Taiqiang, et al.
Published: (2024) -
LLM-NEO: Parameter Efficient Knowledge Distillation for Large Language Models
by: Yang, Runming, et al.
Published: (2024) -
Edge-free but Structure-aware: Prototype-Guided Knowledge Distillation from GNNs to MLPs
by: Wu, Taiqiang, et al.
Published: (2023) -
TS-Memory: Plug-and-Play Memory for Time Series Foundation Models
by: Lyu, Sisuo, et al.
Published: (2026) -
FastQuery: Communication-efficient Embedding Table Query for Private LLM Inference
by: Lin, Chenqi, et al.
Published: (2024)