Early Exit Is a Natural Capability in Transformer-based Models: An Empirical Study on Early Exit without Joint Optimization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shan, Weiqiao, Meng, Long, Zheng, Tong, Luo, Yingfeng, Li, Bei, Wang, junxin, Xiao, Tong, Zhu, Jingbo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Dynamic Early Exit in Reasoning Models
von: Yang, Chenxu, et al.
Veröffentlicht: (2025)
von: Yang, Chenxu, et al.
Veröffentlicht: (2025)
PartialFormer: Modeling Part Instead of Whole for Machine Translation
von: Zheng, Tong, et al.
Veröffentlicht: (2023)
von: Zheng, Tong, et al.
Veröffentlicht: (2023)
BitSkip: An Empirical Analysis of Quantization and Early Exit Composition in Transformers
von: Bhuvaneswaran, Ramshankar, et al.
Veröffentlicht: (2025)
von: Bhuvaneswaran, Ramshankar, et al.
Veröffentlicht: (2025)
FlashThink: An Early Exit Method For Efficient Reasoning
von: Jiang, Guochao, et al.
Veröffentlicht: (2025)
von: Jiang, Guochao, et al.
Veröffentlicht: (2025)
ADEPT: Adaptive Dynamic Early-Exit Process for Transformers
von: Yoo, Sangmin, et al.
Veröffentlicht: (2026)
von: Yoo, Sangmin, et al.
Veröffentlicht: (2026)
EIT: Enhanced Interactive Transformer
von: Zheng, Tong, et al.
Veröffentlicht: (2022)
von: Zheng, Tong, et al.
Veröffentlicht: (2022)
The Zero-Step Thinking: An Empirical Study of Mode Selection as Harder Early Exit in Reasoning Models
von: Tan, Yuqiao, et al.
Veröffentlicht: (2025)
von: Tan, Yuqiao, et al.
Veröffentlicht: (2025)
BEExformer: A Fast Inferencing Binarized Transformer with Early Exits
von: Ansar, Wazib, et al.
Veröffentlicht: (2024)
von: Ansar, Wazib, et al.
Veröffentlicht: (2024)
BEEM: Boosting Performance of Early Exit DNNs using Multi-Exit Classifiers as Experts
von: Bajpai, Divya Jyoti, et al.
Veröffentlicht: (2025)
von: Bajpai, Divya Jyoti, et al.
Veröffentlicht: (2025)
Early-Exit and Instant Confidence Translation Quality Estimation
von: Zouhar, Vilém, et al.
Veröffentlicht: (2025)
von: Zouhar, Vilém, et al.
Veröffentlicht: (2025)
Dynamic Vocabulary Pruning in Early-Exit LLMs
von: Vincenti, Jort, et al.
Veröffentlicht: (2024)
von: Vincenti, Jort, et al.
Veröffentlicht: (2024)
The Diminishing Returns of Early-Exit Decoding in Modern LLMs
von: Wei, Rui, et al.
Veröffentlicht: (2026)
von: Wei, Rui, et al.
Veröffentlicht: (2026)
CEEBERT: Cross-Domain Inference in Early Exit BERT
von: Bajpai, Divya Jyoti, et al.
Veröffentlicht: (2024)
von: Bajpai, Divya Jyoti, et al.
Veröffentlicht: (2024)
One Jump Is All You Need: Short-Cutting Transformers for Early Exit Prediction with One Jump to Fit All Exit Levels
von: Seshadri, Amrit Diggavi
Veröffentlicht: (2025)
von: Seshadri, Amrit Diggavi
Veröffentlicht: (2025)
NEAT: Neuron-Based Early Exit for Large Reasoning Models
von: Liu, Kang, et al.
Veröffentlicht: (2026)
von: Liu, Kang, et al.
Veröffentlicht: (2026)
CAPEEN: Image Captioning with Early Exits and Knowledge Distillation
von: Bajpai, Divya Jyoti, et al.
Veröffentlicht: (2024)
von: Bajpai, Divya Jyoti, et al.
Veröffentlicht: (2024)
DAdEE: Unsupervised Domain Adaptation in Early Exit PLMs
von: Bajpai, Divya Jyoti, et al.
Veröffentlicht: (2024)
von: Bajpai, Divya Jyoti, et al.
Veröffentlicht: (2024)
A Survey of Early Exit Deep Neural Networks in NLP
von: Bajpai, Divya Jyoti, et al.
Veröffentlicht: (2025)
von: Bajpai, Divya Jyoti, et al.
Veröffentlicht: (2025)
Accelerating Large Language Model Inference with Self-Supervised Early Exits
von: Valade, Florian
Veröffentlicht: (2024)
von: Valade, Florian
Veröffentlicht: (2024)
HiDrop: Hierarchical Vision Token Reduction in MLLMs via Late Injection, Concave Pyramid Pruning, and Early Exit
von: Wu, Hao, et al.
Veröffentlicht: (2026)
von: Wu, Hao, et al.
Veröffentlicht: (2026)
Pipeline Parallelism is All You Need for Optimized Early-Exit Based Self-Speculative Decoding
von: Li, Ruanjun, et al.
Veröffentlicht: (2025)
von: Li, Ruanjun, et al.
Veröffentlicht: (2025)
Stop When Reasoning Converges: Semantic-Preserving Early Exit for Reasoning Models
von: Min, Dehai, et al.
Veröffentlicht: (2026)
von: Min, Dehai, et al.
Veröffentlicht: (2026)
RAEE: A Robust Retrieval-Augmented Early Exit Framework for Efficient Inference
von: Huang, Lianming, et al.
Veröffentlicht: (2024)
von: Huang, Lianming, et al.
Veröffentlicht: (2024)
When Is Thinking Enough? Early Exit via Sufficiency Assessment for Efficient Reasoning
von: Xiang, Yang, et al.
Veröffentlicht: (2026)
von: Xiang, Yang, et al.
Veröffentlicht: (2026)
Dissecting Long-Chain-of-Thought Reasoning Models: An Empirical Study
von: Mu, Yongyu, et al.
Veröffentlicht: (2025)
von: Mu, Yongyu, et al.
Veröffentlicht: (2025)
Optimizing Speech Multi-View Feature Fusion through Conditional Computation
von: Shan, Weiqiao, et al.
Veröffentlicht: (2025)
von: Shan, Weiqiao, et al.
Veröffentlicht: (2025)
LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding
von: Elhoushi, Mostafa, et al.
Veröffentlicht: (2024)
von: Elhoushi, Mostafa, et al.
Veröffentlicht: (2024)
HELIOS: Adaptive Model And Early-Exit Selection for Efficient LLM Inference Serving
von: Kumar, Avinash, et al.
Veröffentlicht: (2025)
von: Kumar, Avinash, et al.
Veröffentlicht: (2025)
SPAR-K: Scheduled Periodic Alternating Early Exit for Spoken Language Models
von: Huang, Hsiao-Ying, et al.
Veröffentlicht: (2026)
von: Huang, Hsiao-Ying, et al.
Veröffentlicht: (2026)
SpecExit: Accelerating Large Reasoning Model via Speculative Exit
von: Yang, Rubing, et al.
Veröffentlicht: (2025)
von: Yang, Rubing, et al.
Veröffentlicht: (2025)
Beyond Decoder-only: Large Language Models Can be Good Encoders for Machine Translation
von: Luo, Yingfeng, et al.
Veröffentlicht: (2025)
von: Luo, Yingfeng, et al.
Veröffentlicht: (2025)
Defending against Jailbreak through Early Exit Generation of Large Language Models
von: Zhao, Chongwen, et al.
Veröffentlicht: (2024)
von: Zhao, Chongwen, et al.
Veröffentlicht: (2024)
TERMINATOR: Learning Optimal Exit Points for Early Stopping in Chain-of-Thought Reasoning
von: Nagle, Alliot, et al.
Veröffentlicht: (2026)
von: Nagle, Alliot, et al.
Veröffentlicht: (2026)
SimLens for Early Exit in Large Language Models: Eliciting Accurate Latent Predictions with One More Token
von: Ma, Ming, et al.
Veröffentlicht: (2025)
von: Ma, Ming, et al.
Veröffentlicht: (2025)
Runaway is Ashamed, But Helpful: On the Early-Exit Behavior of Large Language Model-based Agents in Embodied Environments
von: Lu, Qingyu, et al.
Veröffentlicht: (2025)
von: Lu, Qingyu, et al.
Veröffentlicht: (2025)
TIDE: Token-Informed Depth Execution for Per-Token Early Exit in LLM Inference
von: Jaber, Jaber, et al.
Veröffentlicht: (2026)
von: Jaber, Jaber, et al.
Veröffentlicht: (2026)
Federating Dynamic Models using Early-Exit Architectures for Automatic Speech Recognition on Heterogeneous Clients
von: Ali, Mohamed Nabih, et al.
Veröffentlicht: (2024)
von: Ali, Mohamed Nabih, et al.
Veröffentlicht: (2024)
RouteLMT: Learned Sample Routing for Hybrid LLM Translation Deployment
von: Luo, Yingfeng, et al.
Veröffentlicht: (2026)
von: Luo, Yingfeng, et al.
Veröffentlicht: (2026)
Efficient Prompting Methods for Large Language Models: A Survey
von: Chang, Kaiyan, et al.
Veröffentlicht: (2024)
von: Chang, Kaiyan, et al.
Veröffentlicht: (2024)
NiuTrans.LMT: Toward Inclusive and Scalable Multilingual Machine Translation with LLMs
von: Luo, Yingfeng, et al.
Veröffentlicht: (2025)
von: Luo, Yingfeng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Dynamic Early Exit in Reasoning Models
von: Yang, Chenxu, et al.
Veröffentlicht: (2025) -
PartialFormer: Modeling Part Instead of Whole for Machine Translation
von: Zheng, Tong, et al.
Veröffentlicht: (2023) -
BitSkip: An Empirical Analysis of Quantization and Early Exit Composition in Transformers
von: Bhuvaneswaran, Ramshankar, et al.
Veröffentlicht: (2025) -
FlashThink: An Early Exit Method For Efficient Reasoning
von: Jiang, Guochao, et al.
Veröffentlicht: (2025) -
ADEPT: Adaptive Dynamic Early-Exit Process for Transformers
von: Yoo, Sangmin, et al.
Veröffentlicht: (2026)