LEAP: Layer-wise Exit-Aware Pretraining for Efficient Transformer Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Kapadia, Shashank, Mishra, Deep Naryan, Alugubelli, Sujal Reddy, Wang, Haoan, Vabbilisetty, Saipraveen, Bhatia, Rishi, Sharma, Anupriya |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SURGE: SuperBatch Unified Resource-efficient GPU Encoding for Heterogeneous Partitioned Data
by: Kapadia, Shashank, et al.
Published: (2026)
by: Kapadia, Shashank, et al.
Published: (2026)
HELIOS: Adaptive Model And Early-Exit Selection for Efficient LLM Inference Serving
by: Kumar, Avinash, et al.
Published: (2025)
by: Kumar, Avinash, et al.
Published: (2025)
DEL: Context-Aware Dynamic Exit Layer for Efficient Self-Speculative Decoding
by: Zarch, Hossein Entezari, et al.
Published: (2025)
by: Zarch, Hossein Entezari, et al.
Published: (2025)
On the Effect of Uncertainty on Layer-wise Inference Dynamics
by: Kim, Sunwoo, et al.
Published: (2025)
by: Kim, Sunwoo, et al.
Published: (2025)
LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding
by: Elhoushi, Mostafa, et al.
Published: (2024)
by: Elhoushi, Mostafa, et al.
Published: (2024)
BEExformer: A Fast Inferencing Binarized Transformer with Early Exits
by: Ansar, Wazib, et al.
Published: (2024)
by: Ansar, Wazib, et al.
Published: (2024)
RAEE: A Robust Retrieval-Augmented Early Exit Framework for Efficient Inference
by: Huang, Lianming, et al.
Published: (2024)
by: Huang, Lianming, et al.
Published: (2024)
Impact of Layer Norm on Memorization and Generalization in Transformers
by: Singhal, Rishi, et al.
Published: (2025)
by: Singhal, Rishi, et al.
Published: (2025)
LAET: A Layer-wise Adaptive Ensemble Tuning Framework for Pretrained Language Models
by: Ahad, Jawad Ibn, et al.
Published: (2025)
by: Ahad, Jawad Ibn, et al.
Published: (2025)
Suppressing Final Layer Hidden State Jumps in Transformer Pretraining
by: Shibata, Keigo, et al.
Published: (2026)
by: Shibata, Keigo, et al.
Published: (2026)
PromptDistill: Query-based Selective Token Retention in Intermediate Layers for Efficient Large Language Model Inference
by: Jin, Weisheng, et al.
Published: (2025)
by: Jin, Weisheng, et al.
Published: (2025)
Legal Challenges in Regulating Cryptocurrency in India
by: Sujal Sharma
Published: (2026)
by: Sujal Sharma
Published: (2026)
CEEBERT: Cross-Domain Inference in Early Exit BERT
by: Bajpai, Divya Jyoti, et al.
Published: (2024)
by: Bajpai, Divya Jyoti, et al.
Published: (2024)
Efficient Layer-wise LLM Fine-tuning for Revision Intention Prediction
by: Liu, Zhexiong, et al.
Published: (2025)
by: Liu, Zhexiong, et al.
Published: (2025)
Iterative Layer-wise Distillation for Efficient Compression of Large Language Models
by: Kovalev, Grigory, et al.
Published: (2025)
by: Kovalev, Grigory, et al.
Published: (2025)
Relaxed Recursive Transformers: Effective Parameter Sharing with Layer-wise LoRA
by: Bae, Sangmin, et al.
Published: (2024)
by: Bae, Sangmin, et al.
Published: (2024)
Growing Transformers: Modular Composition and Layer-wise Expansion on a Frozen Substrate
by: Bochkov, A.
Published: (2025)
by: Bochkov, A.
Published: (2025)
ELO: Efficient Layer-Specific Optimization for Continual Pretraining of Multilingual LLMs
by: Yoo, HanGyeol, et al.
Published: (2026)
by: Yoo, HanGyeol, et al.
Published: (2026)
Early Exit Is a Natural Capability in Transformer-based Models: An Empirical Study on Early Exit without Joint Optimization
by: Shan, Weiqiao, et al.
Published: (2024)
by: Shan, Weiqiao, et al.
Published: (2024)
TELL-TALE: Task Efficient LLMs with Task Aware Layer Elimination
by: Naim, Omar, et al.
Published: (2025)
by: Naim, Omar, et al.
Published: (2025)
Less, but Better: Efficient Multilingual Expansion for LLMs via Layer-wise Mixture-of-Experts
by: Zhang, Xue, et al.
Published: (2025)
by: Zhang, Xue, et al.
Published: (2025)
Layer-wise Swapping for Generalizable Multilingual Safety
by: Shin, Hyunseo, et al.
Published: (2026)
by: Shin, Hyunseo, et al.
Published: (2026)
Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context Length
by: Ma, Xuezhe, et al.
Published: (2024)
by: Ma, Xuezhe, et al.
Published: (2024)
ADEPT: Adaptive Dynamic Early-Exit Process for Transformers
by: Yoo, Sangmin, et al.
Published: (2026)
by: Yoo, Sangmin, et al.
Published: (2026)
Layered Insights: Generalizable Analysis of Authorial Style by Leveraging All Transformer Layers
by: Alshomary, Milad, et al.
Published: (2025)
by: Alshomary, Milad, et al.
Published: (2025)
Accelerating Large Language Model Inference with Self-Supervised Early Exits
by: Valade, Florian
Published: (2024)
by: Valade, Florian
Published: (2024)
Bridging the Dimensional Chasm: Uncover Layer-wise Dimensional Reduction in Transformers through Token Correlation
by: Song, Zhuo-Yang, et al.
Published: (2025)
by: Song, Zhuo-Yang, et al.
Published: (2025)
DASH: Input-Aware Dynamic Layer Skipping for Efficient LLM Inference with Markov Decision Policies
by: Yang, Ning, et al.
Published: (2025)
by: Yang, Ning, et al.
Published: (2025)
FlashThink: An Early Exit Method For Efficient Reasoning
by: Jiang, Guochao, et al.
Published: (2025)
by: Jiang, Guochao, et al.
Published: (2025)
SqueezeAttention: 2D Management of KV-Cache in LLM Inference via Layer-wise Optimal Budget
by: Wang, Zihao, et al.
Published: (2024)
by: Wang, Zihao, et al.
Published: (2024)
Iterative Layer Pruning for Efficient Translation Inference
by: Moslem, Yasmin, et al.
Published: (2025)
by: Moslem, Yasmin, et al.
Published: (2025)
APT: Adaptive Pruning and Tuning Pretrained Language Models for Efficient Training and Inference
by: Zhao, Bowen, et al.
Published: (2024)
by: Zhao, Bowen, et al.
Published: (2024)
BIPEFT: Budget-Guided Iterative Search for Parameter Efficient Fine-Tuning of Large Pretrained Language Models
by: Chang, Aofei, et al.
Published: (2024)
by: Chang, Aofei, et al.
Published: (2024)
Layer-wise Regularized Dropout for Neural Language Models
by: Ni, Shiwen, et al.
Published: (2024)
by: Ni, Shiwen, et al.
Published: (2024)
Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision
by: Pala, Tej Deep, et al.
Published: (2025)
by: Pala, Tej Deep, et al.
Published: (2025)
LayerBoost: Layer-Aware Attention Reduction for Efficient LLMs
by: Souibgui, Mohamed Ali, et al.
Published: (2026)
by: Souibgui, Mohamed Ali, et al.
Published: (2026)
Dynamic Rebatching for Efficient Early-Exit Inference with DREX
by: Liu, Xuting, et al.
Published: (2025)
by: Liu, Xuting, et al.
Published: (2025)
TIDE: Token-Informed Depth Execution for Per-Token Early Exit in LLM Inference
by: Jaber, Jaber, et al.
Published: (2026)
by: Jaber, Jaber, et al.
Published: (2026)
When Is Thinking Enough? Early Exit via Sufficiency Assessment for Efficient Reasoning
by: Xiang, Yang, et al.
Published: (2026)
by: Xiang, Yang, et al.
Published: (2026)
LaDiMo: Layer-wise Distillation Inspired MoEfier
by: Kim, Sungyoon, et al.
Published: (2024)
by: Kim, Sungyoon, et al.
Published: (2024)
Similar Items
-
SURGE: SuperBatch Unified Resource-efficient GPU Encoding for Heterogeneous Partitioned Data
by: Kapadia, Shashank, et al.
Published: (2026) -
HELIOS: Adaptive Model And Early-Exit Selection for Efficient LLM Inference Serving
by: Kumar, Avinash, et al.
Published: (2025) -
DEL: Context-Aware Dynamic Exit Layer for Efficient Self-Speculative Decoding
by: Zarch, Hossein Entezari, et al.
Published: (2025) -
On the Effect of Uncertainty on Layer-wise Inference Dynamics
by: Kim, Sunwoo, et al.
Published: (2025) -
LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding
by: Elhoushi, Mostafa, et al.
Published: (2024)