EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Xing, Xingrun, Liu, Zheng, Xiao, Shitao, Gao, Boyan, Liang, Yiming, Zhang, Wanpeng, Lin, Haokun, Li, Guoqi, Zhang, Jiajun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SpikeLLM: Scaling up Spiking Neural Network to Large Language Models via Saliency-based Spiking
by: Xing, Xingrun, et al.
Published: (2024)
by: Xing, Xingrun, et al.
Published: (2024)
PretrainZero: Reinforcement Active Pretraining
by: Xing, Xingrun, et al.
Published: (2025)
by: Xing, Xingrun, et al.
Published: (2025)
SpikeLM: Towards General Spike-Driven Language Modeling via Elastic Bi-Spiking Mechanisms
by: Xing, Xingrun, et al.
Published: (2024)
by: Xing, Xingrun, et al.
Published: (2024)
EfficientLLM: Efficiency in Large Language Models
by: Yuan, Zhengqing, et al.
Published: (2025)
by: Yuan, Zhengqing, et al.
Published: (2025)
KLoB: a Benchmark for Assessing Knowledge Locating Methods in Language Models
by: Ju, Yiming, et al.
Published: (2023)
by: Ju, Yiming, et al.
Published: (2023)
Trust Region On-Policy Distillation
by: Xing, Xingrun, et al.
Published: (2026)
by: Xing, Xingrun, et al.
Published: (2026)
From Pixels to Tokens: Byte-Pair Encoding on Quantized Visual Modalities
by: Zhang, Wanpeng, et al.
Published: (2024)
by: Zhang, Wanpeng, et al.
Published: (2024)
Mitigating Training Imbalance in LLM Fine-Tuning via Selective Parameter Merging
by: Ju, Yiming, et al.
Published: (2024)
by: Ju, Yiming, et al.
Published: (2024)
Enhancing Generalization via Sharpness-Aware Trajectory Matching for Dataset Condensation
by: Gao, Boyan, et al.
Published: (2025)
by: Gao, Boyan, et al.
Published: (2025)
Joint-Aligned Latent Action: Towards Scalable VLA Pretraining in the Wild
by: Luo, Hao, et al.
Published: (2026)
by: Luo, Hao, et al.
Published: (2026)
PDTrim: Targeted Pruning for Prefill-Decode Disaggregation in Inference
by: Zhang, Hao, et al.
Published: (2025)
by: Zhang, Hao, et al.
Published: (2025)
Spatial-Aware VLA Pretraining through Visual-Physical Alignment from Human Videos
by: Feng, Yicheng, et al.
Published: (2025)
by: Feng, Yicheng, et al.
Published: (2025)
BESA: Pruning Large Language Models with Blockwise Parameter-Efficient Sparsity Allocation
by: Xu, Peng, et al.
Published: (2024)
by: Xu, Peng, et al.
Published: (2024)
BiPFT: Binary Pre-trained Foundation Transformer with Low-rank Estimation of Binarization Residual Polynomials
by: Xing, Xingrun, et al.
Published: (2023)
by: Xing, Xingrun, et al.
Published: (2023)
Extensible Embedding: A Flexible Multipler For LLM's Context Length
by: Shao, Ninglu, et al.
Published: (2024)
by: Shao, Ninglu, et al.
Published: (2024)
CipherPrune: Efficient and Scalable Private Transformer Inference
by: Zhang, Yancheng, et al.
Published: (2025)
by: Zhang, Yancheng, et al.
Published: (2025)
Spike-driven Large Language Model
by: Xu, Han, et al.
Published: (2026)
by: Xu, Han, et al.
Published: (2026)
AdaRefiner: Refining Decisions of Language Models with Adaptive Feedback
by: Zhang, Wanpeng, et al.
Published: (2023)
by: Zhang, Wanpeng, et al.
Published: (2023)
CachePrune: Privacy-Aware and Fine-Grained KV Cache Sharing for Efficient LLM Inference
by: Wu, Guanlong, et al.
Published: (2026)
by: Wu, Guanlong, et al.
Published: (2026)
FedNano: Toward Lightweight Federated Tuning for Pretrained Multimodal Large Language Models
by: Zhang, Yao, et al.
Published: (2025)
by: Zhang, Yao, et al.
Published: (2025)
Flexibly Scaling Large Language Models Contexts Through Extensible Tokenization
by: Shao, Ninglu, et al.
Published: (2024)
by: Shao, Ninglu, et al.
Published: (2024)
Modality-Aware Zero-Shot Pruning and Sparse Attention for Efficient Multimodal Edge Inference
by: Sui, Yueyuan, et al.
Published: (2026)
by: Sui, Yueyuan, et al.
Published: (2026)
FGP: Feature-Gradient-Prune for Efficient Convolutional Layer Pruning
by: Lv, Qingsong, et al.
Published: (2024)
by: Lv, Qingsong, et al.
Published: (2024)
Continual Learning at the Edge: An Agnostic IIoT Architecture
by: García-Santaclara, Pablo, et al.
Published: (2025)
by: García-Santaclara, Pablo, et al.
Published: (2025)
OmniGen: Unified Image Generation
by: Xiao, Shitao, et al.
Published: (2024)
by: Xiao, Shitao, et al.
Published: (2024)
Dynamic Probabilistic Scheduling for Efficient Edge Resource Orchestration in Tier‐Agnostic Architecture
by: Jingjing Zhu, et al.
Published: (2025)
by: Jingjing Zhu, et al.
Published: (2025)
Environment-Aware Dynamic Pruning for Pipelined Edge Inference
by: O'Quinn, Austin, et al.
Published: (2025)
by: O'Quinn, Austin, et al.
Published: (2025)
LoopLynx: A Scalable Dataflow Architecture for Efficient LLM Inference
by: Zheng, Jianing, et al.
Published: (2025)
by: Zheng, Jianing, et al.
Published: (2025)
CodePMP: Scalable Preference Model Pretraining for Large Language Model Reasoning
by: Yu, Huimu, et al.
Published: (2024)
by: Yu, Huimu, et al.
Published: (2024)
Position-Aware Depth Decay Decoding ($D^3$): Boosting Large Language Model Inference Efficiency
by: Fan, Siqi, et al.
Published: (2025)
by: Fan, Siqi, et al.
Published: (2025)
MoPE-CLIP: Structured Pruning for Efficient Vision-Language Models with Module-wise Pruning Error Metric
by: Lin, Haokun, et al.
Published: (2024)
by: Lin, Haokun, et al.
Published: (2024)
ALPS: Attention Localization and Pruning Strategy for Efficient Alignment of Large Language Models
by: Chen, Hao, et al.
Published: (2025)
by: Chen, Hao, et al.
Published: (2025)
MedREK: Retrieval-Based Editing for Medical LLMs with Key-Aware Prompts
by: Xia, Shujun, et al.
Published: (2025)
by: Xia, Shujun, et al.
Published: (2025)
DPPA: Pruning Method for Large Language Model to Model Merging
by: Zhu, Yaochen, et al.
Published: (2024)
by: Zhu, Yaochen, et al.
Published: (2024)
Diversifying the Expert Knowledge for Task-Agnostic Pruning in Sparse Mixture-of-Experts
by: Zhang, Zeliang, et al.
Published: (2024)
by: Zhang, Zeliang, et al.
Published: (2024)
Magnitude Pruning of Large Pretrained Transformer Models with a Mixture Gaussian Prior
by: Zhang, Mingxuan, et al.
Published: (2024)
by: Zhang, Mingxuan, et al.
Published: (2024)
TENET: An Efficient Sparsity-Aware LUT-Centric Architecture for Ternary LLM Inference On Edge
by: Huang, Zhirui, et al.
Published: (2025)
by: Huang, Zhirui, et al.
Published: (2025)
MINI-LLM: Memory-Efficient Structured Pruning for Large Language Models
by: Cheng, Hongrong, et al.
Published: (2024)
by: Cheng, Hongrong, et al.
Published: (2024)
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
by: Luo, Hao, et al.
Published: (2025)
by: Luo, Hao, et al.
Published: (2025)
APT: Adaptive Pruning and Tuning Pretrained Language Models for Efficient Training and Inference
by: Zhao, Bowen, et al.
Published: (2024)
by: Zhao, Bowen, et al.
Published: (2024)
Similar Items
-
SpikeLLM: Scaling up Spiking Neural Network to Large Language Models via Saliency-based Spiking
by: Xing, Xingrun, et al.
Published: (2024) -
PretrainZero: Reinforcement Active Pretraining
by: Xing, Xingrun, et al.
Published: (2025) -
SpikeLM: Towards General Spike-Driven Language Modeling via Elastic Bi-Spiking Mechanisms
by: Xing, Xingrun, et al.
Published: (2024) -
EfficientLLM: Efficiency in Large Language Models
by: Yuan, Zhengqing, et al.
Published: (2025) -
KLoB: a Benchmark for Assessing Knowledge Locating Methods in Language Models
by: Ju, Yiming, et al.
Published: (2023)