LayerBoost: Layer-Aware Attention Reduction for Efficient LLMs
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Souibgui, Mohamed Ali, Fostier, Jan, Abadía-Heredia, Rodrigo, Denysenko, Bohdan, Marschke, Christian, Peric, Igor |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
TELL-TALE: Task Efficient LLMs with Task Aware Layer Elimination
par: Naim, Omar, et autres
Publié: (2025)
par: Naim, Omar, et autres
Publié: (2025)
DELTA: Dynamic Layer-Aware Token Attention for Efficient Long-Context Reasoning
par: Zarch, Hossein Entezari, et autres
Publié: (2025)
par: Zarch, Hossein Entezari, et autres
Publié: (2025)
Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference
par: Qiu, Quantong, et autres
Publié: (2026)
par: Qiu, Quantong, et autres
Publié: (2026)
High-Layer Attention Pruning with Rescaling
par: Liu, Songtao, et autres
Publié: (2025)
par: Liu, Songtao, et autres
Publié: (2025)
Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse
par: Fu, Zizhuo, et autres
Publié: (2026)
par: Fu, Zizhuo, et autres
Publié: (2026)
Paying Attention to Facts: Quantifying the Knowledge Capacity of Attention Layers
par: Wong, Liang Ze
Publié: (2025)
par: Wong, Liang Ze
Publié: (2025)
Multi-Layer Attention is the Amplifier of Demonstration Effectiveness
par: Wang, Dingzirui, et autres
Publié: (2025)
par: Wang, Dingzirui, et autres
Publié: (2025)
Depth-Wise Attention (DWAtt): A Layer Fusion Method for Data-Efficient Classification
par: ElNokrashy, Muhammad, et autres
Publié: (2022)
par: ElNokrashy, Muhammad, et autres
Publié: (2022)
DEL: Context-Aware Dynamic Exit Layer for Efficient Self-Speculative Decoding
par: Zarch, Hossein Entezari, et autres
Publié: (2025)
par: Zarch, Hossein Entezari, et autres
Publié: (2025)
DLO: Dynamic Layer Operation for Efficient Vertical Scaling of LLMs
par: Tan, Zhen, et autres
Publié: (2024)
par: Tan, Zhen, et autres
Publié: (2024)
An Ensemble Classification Approach in A Multi-Layered Large Language Model Framework for Disease Prediction
par: Hamdi, Ali, et autres
Publié: (2025)
par: Hamdi, Ali, et autres
Publié: (2025)
When Attention Collapses: How Degenerate Layers in LLMs Enable Smaller, Stronger Models
par: Sanyal, Sunny, et autres
Publié: (2024)
par: Sanyal, Sunny, et autres
Publié: (2024)
LEAP: Layer-wise Exit-Aware Pretraining for Efficient Transformer Inference
par: Kapadia, Shashank, et autres
Publié: (2026)
par: Kapadia, Shashank, et autres
Publié: (2026)
A Semantic-Aware Layer-Freezing Approach to Computation-Efficient Fine-Tuning of Language Models
par: Gu, Jian, et autres
Publié: (2024)
par: Gu, Jian, et autres
Publié: (2024)
DASH: Input-Aware Dynamic Layer Skipping for Efficient LLM Inference with Markov Decision Policies
par: Yang, Ning, et autres
Publié: (2025)
par: Yang, Ning, et autres
Publié: (2025)
CLAA: Cross-Layer Attention Aggregation for Accelerating LLM Prefill
par: McDanel, Bradley, et autres
Publié: (2026)
par: McDanel, Bradley, et autres
Publié: (2026)
Mechanism and Emergence of Stacked Attention Heads in Multi-Layer Transformers
par: Musat, Tiberiu
Publié: (2024)
par: Musat, Tiberiu
Publié: (2024)
Discovering the Gems in Early Layers: Accelerating Long-Context LLMs with 1000x Input Token Reduction
par: Shi, Zhenmei, et autres
Publié: (2024)
par: Shi, Zhenmei, et autres
Publié: (2024)
Rethinking Attention: Exploring Shallow Feed-Forward Neural Networks as an Alternative to Attention Layers in Transformers
par: Bozic, Vukasin, et autres
Publié: (2023)
par: Bozic, Vukasin, et autres
Publié: (2023)
Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
par: Brandon, William, et autres
Publié: (2024)
par: Brandon, William, et autres
Publié: (2024)
Wave-PDE Nets: Trainable Wave-Equation Layers as an Alternative to Attention
par: Vejendla, Harshil
Publié: (2025)
par: Vejendla, Harshil
Publié: (2025)
Dr.LLM: Dynamic Layer Routing in LLMs
par: Heakl, Ahmed, et autres
Publié: (2025)
par: Heakl, Ahmed, et autres
Publié: (2025)
Not All Layers of LLMs Are Necessary During Inference
par: Fan, Siqi, et autres
Publié: (2024)
par: Fan, Siqi, et autres
Publié: (2024)
IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse
par: Bai, Yushi, et autres
Publié: (2026)
par: Bai, Yushi, et autres
Publié: (2026)
AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers
par: Achtibat, Reduan, et autres
Publié: (2024)
par: Achtibat, Reduan, et autres
Publié: (2024)
Layer-wise Representation Dynamics: An Empirical Investigation Across Embedders and Base LLMs
par: Jiang, Jingzhou, et autres
Publié: (2026)
par: Jiang, Jingzhou, et autres
Publié: (2026)
Internal Chain-of-Thought: Empirical Evidence for Layer-wise Subtask Scheduling in LLMs
par: Yang, Zhipeng, et autres
Publié: (2025)
par: Yang, Zhipeng, et autres
Publié: (2025)
A Comparative analysis of Layer-wise Representational Capacity in AR and Diffusion LLMs
par: Goel, Raghavv, et autres
Publié: (2026)
par: Goel, Raghavv, et autres
Publié: (2026)
Calibration Across Layers: Understanding Calibration Evolution in LLMs
par: Joshi, Abhinav, et autres
Publié: (2025)
par: Joshi, Abhinav, et autres
Publié: (2025)
Towards Understanding the Word Sensitivity of Attention Layers: A Study via Random Features
par: Bombari, Simone, et autres
Publié: (2024)
par: Bombari, Simone, et autres
Publié: (2024)
Sparse Query Attention (SQA): A Computationally Efficient Attention Mechanism with Query Heads Reduction
par: Filipek, Adam
Publié: (2025)
par: Filipek, Adam
Publié: (2025)
Bridging the Dimensional Chasm: Uncover Layer-wise Dimensional Reduction in Transformers through Token Correlation
par: Song, Zhuo-Yang, et autres
Publié: (2025)
par: Song, Zhuo-Yang, et autres
Publié: (2025)
Efficient LLM Moderation with Multi-Layer Latent Prototypes
par: Chrabąszcz, Maciej, et autres
Publié: (2025)
par: Chrabąszcz, Maciej, et autres
Publié: (2025)
Layer-Aware Task Arithmetic: Disentangling Task-Specific and Instruction-Following Knowledge
par: Chen, Yan-Lun, et autres
Publié: (2025)
par: Chen, Yan-Lun, et autres
Publié: (2025)
Iterative Layer-wise Distillation for Efficient Compression of Large Language Models
par: Kovalev, Grigory, et autres
Publié: (2025)
par: Kovalev, Grigory, et autres
Publié: (2025)
Towards Building Efficient Sentence BERT Models using Layer Pruning
par: Shelke, Anushka, et autres
Publié: (2024)
par: Shelke, Anushka, et autres
Publié: (2024)
Out-of-Distribution Detection by Leveraging Between-Layer Transformation Smoothness
par: Jelenić, Fran, et autres
Publié: (2023)
par: Jelenić, Fran, et autres
Publié: (2023)
Confidence-Credibility Aware Weighted Ensembles of Small LLMs Outperform Large LLMs in Emotion Detection
par: Elgabry, Menna, et autres
Publié: (2025)
par: Elgabry, Menna, et autres
Publié: (2025)
TrimLLM: Progressive Layer Dropping for Domain-Specific LLMs
par: Hu, Lanxiang, et autres
Publié: (2024)
par: Hu, Lanxiang, et autres
Publié: (2024)
ConfLayers: Adaptive Confidence-based Layer Skipping for Self-Speculative Decoding
par: Amer, Walaa, et autres
Publié: (2026)
par: Amer, Walaa, et autres
Publié: (2026)
Documents similaires
-
TELL-TALE: Task Efficient LLMs with Task Aware Layer Elimination
par: Naim, Omar, et autres
Publié: (2025) -
DELTA: Dynamic Layer-Aware Token Attention for Efficient Long-Context Reasoning
par: Zarch, Hossein Entezari, et autres
Publié: (2025) -
Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference
par: Qiu, Quantong, et autres
Publié: (2026) -
High-Layer Attention Pruning with Rescaling
par: Liu, Songtao, et autres
Publié: (2025) -
Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse
par: Fu, Zizhuo, et autres
Publié: (2026)