The Diminishing Returns of Early-Exit Decoding in Modern LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Wei, Rui, Du, Rui, Yu, Hanfei, Tiwari, Devesh, Li, Jian, Xu, Zhaozhuo, Wang, Hao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Self-Consistency Is Losing Its Edge: Diminishing Returns and Rising Costs in Modern LLMs
by: Loo, Chiyan
Published: (2025)
by: Loo, Chiyan
Published: (2025)
RLHFless: Serverless Computing for Efficient RLHF
by: Wei, Rui, et al.
Published: (2026)
by: Wei, Rui, et al.
Published: (2026)
Outliers and Calibration Sets have Diminishing Effect on Quantization of Modern LLMs
by: Paglieri, Davide, et al.
Published: (2024)
by: Paglieri, Davide, et al.
Published: (2024)
Pipeline Parallelism is All You Need for Optimized Early-Exit Based Self-Speculative Decoding
by: Li, Ruanjun, et al.
Published: (2025)
by: Li, Ruanjun, et al.
Published: (2025)
Dynamic Vocabulary Pruning in Early-Exit LLMs
by: Vincenti, Jort, et al.
Published: (2024)
by: Vincenti, Jort, et al.
Published: (2024)
Dynamic Early Exit in Reasoning Models
by: Yang, Chenxu, et al.
Published: (2025)
by: Yang, Chenxu, et al.
Published: (2025)
LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding
by: Elhoushi, Mostafa, et al.
Published: (2024)
by: Elhoushi, Mostafa, et al.
Published: (2024)
ADEPT: Adaptive Dynamic Early-Exit Process for Transformers
by: Yoo, Sangmin, et al.
Published: (2026)
by: Yoo, Sangmin, et al.
Published: (2026)
FlashThink: An Early Exit Method For Efficient Reasoning
by: Jiang, Guochao, et al.
Published: (2025)
by: Jiang, Guochao, et al.
Published: (2025)
English K_Quantization of LLMs Does Not Disproportionately Diminish Multilingual Performance
by: Borgersen, Karl Audun, et al.
Published: (2025)
by: Borgersen, Karl Audun, et al.
Published: (2025)
DAdEE: Unsupervised Domain Adaptation in Early Exit PLMs
by: Bajpai, Divya Jyoti, et al.
Published: (2024)
by: Bajpai, Divya Jyoti, et al.
Published: (2024)
Sampling-Efficient Test-Time Scaling: Self-Estimating the Best-of-N Sampling in Early Decoding
by: Wang, Yiming, et al.
Published: (2025)
by: Wang, Yiming, et al.
Published: (2025)
Dynamic Depth Decoding: Faster Speculative Decoding for LLMs
by: Brown, Oscar, et al.
Published: (2024)
by: Brown, Oscar, et al.
Published: (2024)
A Hybrid DeBERTa and Gated Broad Learning System for Cyberbullying Detection in English Text
by: Kumar, Devesh
Published: (2025)
by: Kumar, Devesh
Published: (2025)
Every Step Counts: Decoding Trajectories as Authorship Fingerprints of dLLMs
by: Li, Qi, et al.
Published: (2025)
by: Li, Qi, et al.
Published: (2025)
DEL-ToM: Inference-Time Scaling for Theory-of-Mind Reasoning via Dynamic Epistemic Logic
by: Wu, Yuheng, et al.
Published: (2025)
by: Wu, Yuheng, et al.
Published: (2025)
Multi-Personality Generation of LLMs at Decoding-time
by: Chen, Rongxin, et al.
Published: (2025)
by: Chen, Rongxin, et al.
Published: (2025)
Beyond Chain-of-Thought: A Survey of Chain-of-X Paradigms for LLMs
by: Xia, Yu, et al.
Published: (2024)
by: Xia, Yu, et al.
Published: (2024)
BEExformer: A Fast Inferencing Binarized Transformer with Early Exits
by: Ansar, Wazib, et al.
Published: (2024)
by: Ansar, Wazib, et al.
Published: (2024)
CAPEEN: Image Captioning with Early Exits and Knowledge Distillation
by: Bajpai, Divya Jyoti, et al.
Published: (2024)
by: Bajpai, Divya Jyoti, et al.
Published: (2024)
The Zero-Step Thinking: An Empirical Study of Mode Selection as Harder Early Exit in Reasoning Models
by: Tan, Yuqiao, et al.
Published: (2025)
by: Tan, Yuqiao, et al.
Published: (2025)
Return of the Encoder: Maximizing Parameter Efficiency for SLMs
by: Elfeki, Mohamed, et al.
Published: (2025)
by: Elfeki, Mohamed, et al.
Published: (2025)
Mitigating Context-Memory Conflicts in LLMs through Dynamic Cognitive Reconciliation Decoding
by: Zhou, Yigeng, et al.
Published: (2026)
by: Zhou, Yigeng, et al.
Published: (2026)
Hybrid Policy Distillation for LLMs
by: Zhu, Wenhong, et al.
Published: (2026)
by: Zhu, Wenhong, et al.
Published: (2026)
Runaway is Ashamed, But Helpful: On the Early-Exit Behavior of Large Language Model-based Agents in Embodied Environments
by: Lu, Qingyu, et al.
Published: (2025)
by: Lu, Qingyu, et al.
Published: (2025)
Toward Sustainable GenAI using Generation Directives for Carbon-Friendly Large Language Model Inference
by: Li, Baolin, et al.
Published: (2024)
by: Li, Baolin, et al.
Published: (2024)
Batch Speculative Decoding Done Right
by: Zhang, Ranran Haoran, et al.
Published: (2025)
by: Zhang, Ranran Haoran, et al.
Published: (2025)
SpecExit: Accelerating Large Reasoning Model via Speculative Exit
by: Yang, Rubing, et al.
Published: (2025)
by: Yang, Rubing, et al.
Published: (2025)
One Jump Is All You Need: Short-Cutting Transformers for Early Exit Prediction with One Jump to Fit All Exit Levels
by: Seshadri, Amrit Diggavi
Published: (2025)
by: Seshadri, Amrit Diggavi
Published: (2025)
Extended Inductive Reasoning for Personalized Preference Inference from Behavioral Signals
by: Li, Jia-Nan, et al.
Published: (2025)
by: Li, Jia-Nan, et al.
Published: (2025)
Ranking LLMs by compression
by: Guo, Peijia, et al.
Published: (2024)
by: Guo, Peijia, et al.
Published: (2024)
Decoding at the Speed of Thought: Harnessing Parallel Decoding of Lexical Units for LLMs
by: Sun, Chenxi, et al.
Published: (2024)
by: Sun, Chenxi, et al.
Published: (2024)
Draft Model Knows When to Stop: Self-Verification Speculative Decoding for Long-Form Generation
by: Zhang, Ziyin, et al.
Published: (2024)
by: Zhang, Ziyin, et al.
Published: (2024)
EE-Tuning: An Economical yet Scalable Solution for Tuning Early-Exit Large Language Models
by: Pan, Xuchen, et al.
Published: (2024)
by: Pan, Xuchen, et al.
Published: (2024)
Towards Compositional Generalization of LLMs via Skill Taxonomy Guided Data Synthesis
by: Wei, Yifan, et al.
Published: (2026)
by: Wei, Yifan, et al.
Published: (2026)
Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression
by: Hong, Junyuan, et al.
Published: (2024)
by: Hong, Junyuan, et al.
Published: (2024)
ASPD: Unlocking Adaptive Serial-Parallel Decoding by Exploring Intrinsic Parallelism in LLMs
by: Chen, Keyu, et al.
Published: (2025)
by: Chen, Keyu, et al.
Published: (2025)
TERMINATOR: Learning Optimal Exit Points for Early Stopping in Chain-of-Thought Reasoning
by: Nagle, Alliot, et al.
Published: (2026)
by: Nagle, Alliot, et al.
Published: (2026)
Defending against Jailbreak through Early Exit Generation of Large Language Models
by: Zhao, Chongwen, et al.
Published: (2024)
by: Zhao, Chongwen, et al.
Published: (2024)
Automated Clinical Data Extraction with Knowledge Conditioned LLMs
by: Li, Diya, et al.
Published: (2024)
by: Li, Diya, et al.
Published: (2024)
Similar Items
-
Self-Consistency Is Losing Its Edge: Diminishing Returns and Rising Costs in Modern LLMs
by: Loo, Chiyan
Published: (2025) -
RLHFless: Serverless Computing for Efficient RLHF
by: Wei, Rui, et al.
Published: (2026) -
Outliers and Calibration Sets have Diminishing Effect on Quantization of Modern LLMs
by: Paglieri, Davide, et al.
Published: (2024) -
Pipeline Parallelism is All You Need for Optimized Early-Exit Based Self-Speculative Decoding
by: Li, Ruanjun, et al.
Published: (2025) -
Dynamic Vocabulary Pruning in Early-Exit LLMs
by: Vincenti, Jort, et al.
Published: (2024)