DASH: Input-Aware Dynamic Layer Skipping for Efficient LLM Inference with Markov Decision Policies
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Ning, Liu, Fangxin, Wang, Junjie, Yang, Tao, Liu, Kan, Guan, Haibing, Jiang, Li |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Calibrated Speculative Decoding: Frequency-Guided Candidate Selection for Efficient Inference
by: Zhou, Xuwen, et al.
Published: (2026)
by: Zhou, Xuwen, et al.
Published: (2026)
LCD: Advancing Extreme Low-Bit Clustering for Large Language Models via Knowledge Distillation
by: Liu, Fangxin, et al.
Published: (2025)
by: Liu, Fangxin, et al.
Published: (2025)
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking
by: Federici, Marco, et al.
Published: (2024)
by: Federici, Marco, et al.
Published: (2024)
Probe and Skip: Self-Predictive Token Skipping for Efficient Long-Context LLM Inference
by: Wu, Zimeng, et al.
Published: (2026)
by: Wu, Zimeng, et al.
Published: (2026)
Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
by: Tang, Jiaming, et al.
Published: (2024)
by: Tang, Jiaming, et al.
Published: (2024)
SpecQuant: Spectral Decomposition and Adaptive Truncation for Ultra-Low-Bit LLMs Quantization
by: Zhao, Zhixiong, et al.
Published: (2025)
by: Zhao, Zhixiong, et al.
Published: (2025)
FFN-SkipLLM: A Hidden Gem for Autoregressive Decoding with Adaptive Feed Forward Skipping
by: Jaiswal, Ajay, et al.
Published: (2024)
by: Jaiswal, Ajay, et al.
Published: (2024)
Learning to Skip the Middle Layers of Transformers
by: Lawson, Tim, et al.
Published: (2025)
by: Lawson, Tim, et al.
Published: (2025)
LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding
by: Elhoushi, Mostafa, et al.
Published: (2024)
by: Elhoushi, Mostafa, et al.
Published: (2024)
DASH: Fast Differentiable Architecture Search for Hybrid Attention in Minutes on a Single GPU
by: Chen, Weizhe, et al.
Published: (2026)
by: Chen, Weizhe, et al.
Published: (2026)
Skipformer: A Skip-and-Recover Strategy for Efficient Speech Recognition
by: Zhu, Wenjing, et al.
Published: (2024)
by: Zhu, Wenjing, et al.
Published: (2024)
FlexQuant: A Flexible and Efficient Dynamic Precision Switching Framework for LLM Quantization
by: Liu, Fangxin, et al.
Published: (2025)
by: Liu, Fangxin, et al.
Published: (2025)
Skip-Connected Policy Optimization for Implicit Advantage
by: Teng, Fengwei, et al.
Published: (2026)
by: Teng, Fengwei, et al.
Published: (2026)
DELTA: Dynamic Layer-Aware Token Attention for Efficient Long-Context Reasoning
by: Zarch, Hossein Entezari, et al.
Published: (2025)
by: Zarch, Hossein Entezari, et al.
Published: (2025)
DEL: Context-Aware Dynamic Exit Layer for Efficient Self-Speculative Decoding
by: Zarch, Hossein Entezari, et al.
Published: (2025)
by: Zarch, Hossein Entezari, et al.
Published: (2025)
ConfLayers: Adaptive Confidence-based Layer Skipping for Self-Speculative Decoding
by: Amer, Walaa, et al.
Published: (2026)
by: Amer, Walaa, et al.
Published: (2026)
KVTuner: Sensitivity-Aware Layer-Wise Mixed-Precision KV Cache Quantization for Efficient and Nearly Lossless LLM Inference
by: Li, Xing, et al.
Published: (2025)
by: Li, Xing, et al.
Published: (2025)
Spark: Strategic Policy-Aware Exploration via Dynamic Branching for Long-Horizon Agentic Learning
by: Wu, Jinyang, et al.
Published: (2026)
by: Wu, Jinyang, et al.
Published: (2026)
Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference
by: Qiu, Quantong, et al.
Published: (2026)
by: Qiu, Quantong, et al.
Published: (2026)
QUARK: Quantization-Enabled Circuit Sharing for Transformer Acceleration by Exploiting Common Patterns in Nonlinear Operations
by: Zhao, Zhixiong, et al.
Published: (2025)
by: Zhao, Zhixiong, et al.
Published: (2025)
Federated Learning with Layer Skipping: Efficient Training of Large Language Models for Healthcare NLP
by: Zhang, Lihong, et al.
Published: (2025)
by: Zhang, Lihong, et al.
Published: (2025)
Inference-Cost-Aware Dynamic Tree Construction for Efficient Inference in Large Language Models
by: Hong, Yinrong, et al.
Published: (2025)
by: Hong, Yinrong, et al.
Published: (2025)
Hardware-Aware Parallel Prompt Decoding for Memory-Efficient Acceleration of LLM Inference
by: Chen, Hao Mark, et al.
Published: (2024)
by: Chen, Hao Mark, et al.
Published: (2024)
PoSE: Efficient Context Window Extension of LLMs via Positional Skip-wise Training
by: Zhu, Dawei, et al.
Published: (2023)
by: Zhu, Dawei, et al.
Published: (2023)
LEAP: Layer-wise Exit-Aware Pretraining for Efficient Transformer Inference
by: Kapadia, Shashank, et al.
Published: (2026)
by: Kapadia, Shashank, et al.
Published: (2026)
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
by: Liu, Guangda, et al.
Published: (2025)
by: Liu, Guangda, et al.
Published: (2025)
DynSplit-KV: Dynamic Semantic Splitting for KVCache Compression in Efficient Long-Context LLM Inference
by: Ye, Jiancai, et al.
Published: (2026)
by: Ye, Jiancai, et al.
Published: (2026)
Internalize the Temperature: On-Policy Self-Distillation as Policy Reheater for Reinforcement Learning
by: Yang, Xuewei, et al.
Published: (2026)
by: Yang, Xuewei, et al.
Published: (2026)
Reliability-Aware Adaptive Self-Consistency for Efficient Sampling in LLM Reasoning
by: Kim, Junseok, et al.
Published: (2026)
by: Kim, Junseok, et al.
Published: (2026)
KVSharer: Efficient Inference via Layer-Wise Dissimilar KV Cache Sharing
by: Yang, Yifei, et al.
Published: (2024)
by: Yang, Yifei, et al.
Published: (2024)
On the Effect of Uncertainty on Layer-wise Inference Dynamics
by: Kim, Sunwoo, et al.
Published: (2025)
by: Kim, Sunwoo, et al.
Published: (2025)
LayerBoost: Layer-Aware Attention Reduction for Efficient LLMs
by: Souibgui, Mohamed Ali, et al.
Published: (2026)
by: Souibgui, Mohamed Ali, et al.
Published: (2026)
Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context Length
by: Ma, Xuezhe, et al.
Published: (2024)
by: Ma, Xuezhe, et al.
Published: (2024)
LayerRoute: Input-Conditioned Adaptive Layer Skipping via LoRA Fine-Tuning for Agentic Language Models
by: Sikdar, Prateek Kumar
Published: (2026)
by: Sikdar, Prateek Kumar
Published: (2026)
Skip-It? Theoretical Conditions for Layer Skipping in Vision-Language Models
by: Hartman, Max, et al.
Published: (2025)
by: Hartman, Max, et al.
Published: (2025)
ONTO: A Token-Efficient Columnar Notation for LLM Input Optimization
by: Deekeswar, Harshavardhanan
Published: (2026)
by: Deekeswar, Harshavardhanan
Published: (2026)
Layer-wise Representation Dynamics: An Empirical Investigation Across Embedders and Base LLMs
by: Jiang, Jingzhou, et al.
Published: (2026)
by: Jiang, Jingzhou, et al.
Published: (2026)
STEP: Success-Rate-Aware Trajectory-Efficient Policy Optimization
by: Chen, Yuhan, et al.
Published: (2025)
by: Chen, Yuhan, et al.
Published: (2025)
LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
by: Fu, Qichen, et al.
Published: (2024)
by: Fu, Qichen, et al.
Published: (2024)
Attention-Aware GNN-based Input Defense against Multi-Turn LLM Jailbreak
by: Huang, Zixuan, et al.
Published: (2025)
by: Huang, Zixuan, et al.
Published: (2025)
Similar Items
-
Calibrated Speculative Decoding: Frequency-Guided Candidate Selection for Efficient Inference
by: Zhou, Xuwen, et al.
Published: (2026) -
LCD: Advancing Extreme Low-Bit Clustering for Large Language Models via Knowledge Distillation
by: Liu, Fangxin, et al.
Published: (2025) -
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking
by: Federici, Marco, et al.
Published: (2024) -
Probe and Skip: Self-Predictive Token Skipping for Efficient Long-Context LLM Inference
by: Wu, Zimeng, et al.
Published: (2026) -
Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
by: Tang, Jiaming, et al.
Published: (2024)