Learning to Skip the Middle Layers of Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Lawson, Tim, Aitchison, Laurence |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Residual Stream Analysis with Multi-Layer SAEs
by: Lawson, Tim, et al.
Published: (2024)
by: Lawson, Tim, et al.
Published: (2024)
Jacobian Sparse Autoencoders: Sparsify Computations, Not Just Activations
by: Farnik, Lucy, et al.
Published: (2025)
by: Farnik, Lucy, et al.
Published: (2025)
Scale-invariant Attention
by: Anson, Ben, et al.
Published: (2025)
by: Anson, Ben, et al.
Published: (2025)
Automated Interpretability Metrics Do Not Distinguish Trained and Random Transformers
by: Heap, Thomas, et al.
Published: (2025)
by: Heap, Thomas, et al.
Published: (2025)
ConfLayers: Adaptive Confidence-based Layer Skipping for Self-Speculative Decoding
by: Amer, Walaa, et al.
Published: (2026)
by: Amer, Walaa, et al.
Published: (2026)
Questionable practices in machine learning
by: Leech, Gavin, et al.
Published: (2024)
by: Leech, Gavin, et al.
Published: (2024)
Federated Learning with Layer Skipping: Efficient Training of Large Language Models for Healthcare NLP
by: Zhang, Lihong, et al.
Published: (2025)
by: Zhang, Lihong, et al.
Published: (2025)
Skip-It? Theoretical Conditions for Layer Skipping in Vision-Language Models
by: Hartman, Max, et al.
Published: (2025)
by: Hartman, Max, et al.
Published: (2025)
DASH: Input-Aware Dynamic Layer Skipping for Efficient LLM Inference with Markov Decision Policies
by: Yang, Ning, et al.
Published: (2025)
by: Yang, Ning, et al.
Published: (2025)
Probe and Skip: Self-Predictive Token Skipping for Efficient Long-Context LLM Inference
by: Wu, Zimeng, et al.
Published: (2026)
by: Wu, Zimeng, et al.
Published: (2026)
LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding
by: Elhoushi, Mostafa, et al.
Published: (2024)
by: Elhoushi, Mostafa, et al.
Published: (2024)
FFN-SkipLLM: A Hidden Gem for Autoregressive Decoding with Adaptive Feed Forward Skipping
by: Jaiswal, Ajay, et al.
Published: (2024)
by: Jaiswal, Ajay, et al.
Published: (2024)
Skip-Connected Policy Optimization for Implicit Advantage
by: Teng, Fengwei, et al.
Published: (2026)
by: Teng, Fengwei, et al.
Published: (2026)
Position-Agnostic Pre-Projection for Transformer Attention: Nonlinear Feature Construction and Content Skip Before Q/K/V
by: Shinde, Chirag
Published: (2026)
by: Shinde, Chirag
Published: (2026)
Why you don't overfit, and don't need Bayes if you only train for one epoch
by: Aitchison, Laurence
Published: (2024)
by: Aitchison, Laurence
Published: (2024)
Skipformer: A Skip-and-Recover Strategy for Efficient Speech Recognition
by: Zhu, Wenjing, et al.
Published: (2024)
by: Zhu, Wenjing, et al.
Published: (2024)
On SkipGram Word Embedding Models with Negative Sampling: Unified Framework and Impact of Noise Distributions
by: Liu, Dezhi, et al.
Published: (2020)
by: Liu, Dezhi, et al.
Published: (2020)
PoSE: Efficient Context Window Extension of LLMs via Positional Skip-wise Training
by: Zhu, Dawei, et al.
Published: (2023)
by: Zhu, Dawei, et al.
Published: (2023)
LayerNorm Induces Recency Bias in Transformer Decoders
by: Kim, Junu, et al.
Published: (2025)
by: Kim, Junu, et al.
Published: (2025)
Provable Knowledge Acquisition and Extraction in One-Layer Transformers
by: Xu, Ruichen, et al.
Published: (2025)
by: Xu, Ruichen, et al.
Published: (2025)
Scavenging Hyena: Distilling Transformers into Long Convolution Models
by: Ralambomihanta, Tokiniaina Raharison, et al.
Published: (2024)
by: Ralambomihanta, Tokiniaina Raharison, et al.
Published: (2024)
Out-of-Distribution Detection by Leveraging Between-Layer Transformation Smoothness
by: Jelenić, Fran, et al.
Published: (2023)
by: Jelenić, Fran, et al.
Published: (2023)
Mechanism and Emergence of Stacked Attention Heads in Multi-Layer Transformers
by: Musat, Tiberiu
Published: (2024)
by: Musat, Tiberiu
Published: (2024)
Lost in the Middle at Birth: An Exact Theory of Transformer Position Bias
by: Chowdhury, Borun D
Published: (2026)
by: Chowdhury, Borun D
Published: (2026)
Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
by: Brandon, William, et al.
Published: (2024)
by: Brandon, William, et al.
Published: (2024)
Growing Transformers: Modular Composition and Layer-wise Expansion on a Frozen Substrate
by: Bochkov, A.
Published: (2025)
by: Bochkov, A.
Published: (2025)
Relaxed Recursive Transformers: Effective Parameter Sharing with Layer-wise LoRA
by: Bae, Sangmin, et al.
Published: (2024)
by: Bae, Sangmin, et al.
Published: (2024)
Controlling changes to attention logits
by: Anson, Ben, et al.
Published: (2025)
by: Anson, Ben, et al.
Published: (2025)
Batch size invariant Adam
by: Wang, Xi, et al.
Published: (2024)
by: Wang, Xi, et al.
Published: (2024)
Post-Trained MoE Can Skip Half Experts via Self-Distillation
by: Lv, Xingtai, et al.
Published: (2026)
by: Lv, Xingtai, et al.
Published: (2026)
Function-Space Learning Rates
by: Milsom, Edward, et al.
Published: (2025)
by: Milsom, Edward, et al.
Published: (2025)
Bridging the Dimensional Chasm: Uncover Layer-wise Dimensional Reduction in Transformers through Token Correlation
by: Song, Zhuo-Yang, et al.
Published: (2025)
by: Song, Zhuo-Yang, et al.
Published: (2025)
Rethinking Attention: Exploring Shallow Feed-Forward Neural Networks as an Alternative to Attention Layers in Transformers
by: Bozic, Vukasin, et al.
Published: (2023)
by: Bozic, Vukasin, et al.
Published: (2023)
Layer by Layer: Uncovering Where Multi-Task Learning Happens in Instruction-Tuned Large Language Models
by: Zhao, Zheng, et al.
Published: (2024)
by: Zhao, Zheng, et al.
Published: (2024)
ResidualTransformer: Residual Low-Rank Learning with Weight-Sharing for Transformer Layers
by: Wang, Yiming, et al.
Published: (2023)
by: Wang, Yiming, et al.
Published: (2023)
The Truth Lies Somewhere in the Middle (of the Generated Tokens)
by: Wang, Sophie L., et al.
Published: (2026)
by: Wang, Sophie L., et al.
Published: (2026)
Dual Path Attribution: Efficient Attribution for SwiGLU-Transformers through Layer-Wise Target Propagation
by: Jantsch, Lasse Marten, et al.
Published: (2026)
by: Jantsch, Lasse Marten, et al.
Published: (2026)
Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models
by: Lu, Xudong, et al.
Published: (2024)
by: Lu, Xudong, et al.
Published: (2024)
Peri-LN: Revisiting Normalization Layer in the Transformer Architecture
by: Kim, Jeonghoon, et al.
Published: (2025)
by: Kim, Jeonghoon, et al.
Published: (2025)
Stacking Small Language Models for Generalizability
by: Liang, Laurence
Published: (2024)
by: Liang, Laurence
Published: (2024)
Similar Items
-
Residual Stream Analysis with Multi-Layer SAEs
by: Lawson, Tim, et al.
Published: (2024) -
Jacobian Sparse Autoencoders: Sparsify Computations, Not Just Activations
by: Farnik, Lucy, et al.
Published: (2025) -
Scale-invariant Attention
by: Anson, Ben, et al.
Published: (2025) -
Automated Interpretability Metrics Do Not Distinguish Trained and Random Transformers
by: Heap, Thomas, et al.
Published: (2025) -
ConfLayers: Adaptive Confidence-based Layer Skipping for Self-Speculative Decoding
by: Amer, Walaa, et al.
Published: (2026)