High-Layer Attention Pruning with Rescaling
Fuente:
arXiv
Guardado en:
| Autores principales: | Liu, Songtao, Liu, Peng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Condense, Don't Just Prune: Enhancing Efficiency and Performance in MoE Layer Pruning
por: Cao, Mingyu, et al.
Publicado: (2024)
por: Cao, Mingyu, et al.
Publicado: (2024)
Variance Control via Weight Rescaling in LLM Pre-training
por: Owen, Louis, et al.
Publicado: (2025)
por: Owen, Louis, et al.
Publicado: (2025)
SAP: Syntactic Attention Pruning for Transformer-based Language Models
por: Lee, Tzu-Yun, et al.
Publicado: (2025)
por: Lee, Tzu-Yun, et al.
Publicado: (2025)
Twilight: Adaptive Attention Sparsity with Hierarchical Top-$p$ Pruning
por: Lin, Chaofan, et al.
Publicado: (2025)
por: Lin, Chaofan, et al.
Publicado: (2025)
STAR: Spectral Truncation and Rescale for Model Merging
por: Lee, Yu-Ang, et al.
Publicado: (2025)
por: Lee, Yu-Ang, et al.
Publicado: (2025)
On Importance of Layer Pruning for Smaller BERT Models and Low Resource Languages
por: Shirke, Mayur, et al.
Publicado: (2025)
por: Shirke, Mayur, et al.
Publicado: (2025)
Towards Building Efficient Sentence BERT Models using Layer Pruning
por: Shelke, Anushka, et al.
Publicado: (2024)
por: Shelke, Anushka, et al.
Publicado: (2024)
MaskPrune: Mask-based LLM Pruning for Layer-wise Uniform Structures
por: Qin, Jiayu, et al.
Publicado: (2025)
por: Qin, Jiayu, et al.
Publicado: (2025)
LayerBoost: Layer-Aware Attention Reduction for Efficient LLMs
por: Souibgui, Mohamed Ali, et al.
Publicado: (2026)
por: Souibgui, Mohamed Ali, et al.
Publicado: (2026)
Paying Attention to Facts: Quantifying the Knowledge Capacity of Attention Layers
por: Wong, Liang Ze
Publicado: (2025)
por: Wong, Liang Ze
Publicado: (2025)
Adaptive Layer Selection for Layer-Wise Token Pruning in LLM Inference
por: Taniguchi, Rei, et al.
Publicado: (2026)
por: Taniguchi, Rei, et al.
Publicado: (2026)
Tackling Length Inflation Without Trade-offs: Group Relative Reward Rescaling for Reinforcement Learning
por: Li, Zichao, et al.
Publicado: (2026)
por: Li, Zichao, et al.
Publicado: (2026)
Domain-Specific Pruning of Large Mixture-of-Experts Models with Few-shot Demonstrations
por: Dong, Zican, et al.
Publicado: (2025)
por: Dong, Zican, et al.
Publicado: (2025)
Multi-Layer Attention is the Amplifier of Demonstration Effectiveness
por: Wang, Dingzirui, et al.
Publicado: (2025)
por: Wang, Dingzirui, et al.
Publicado: (2025)
Pruning Literals for Highly Efficient Explainability at Word Level
por: Yadav, Rohan Kumar, et al.
Publicado: (2024)
por: Yadav, Rohan Kumar, et al.
Publicado: (2024)
DenoiseRotator: Enhance Pruning Robustness for LLMs via Importance Concentration
por: Gu, Tianteng, et al.
Publicado: (2025)
por: Gu, Tianteng, et al.
Publicado: (2025)
A2SF: Accumulative Attention Scoring with Forgetting Factor for Token Pruning in Transformer Decoder
por: Jo, Hyun-rae, et al.
Publicado: (2024)
por: Jo, Hyun-rae, et al.
Publicado: (2024)
Rethinking Attention: Exploring Shallow Feed-Forward Neural Networks as an Alternative to Attention Layers in Transformers
por: Bozic, Vukasin, et al.
Publicado: (2023)
por: Bozic, Vukasin, et al.
Publicado: (2023)
Frustratingly Easy Task-aware Pruning for Large Language Models
por: Tian, Yuanhe, et al.
Publicado: (2025)
por: Tian, Yuanhe, et al.
Publicado: (2025)
CLAA: Cross-Layer Attention Aggregation for Accelerating LLM Prefill
por: McDanel, Bradley, et al.
Publicado: (2026)
por: McDanel, Bradley, et al.
Publicado: (2026)
Mechanism and Emergence of Stacked Attention Heads in Multi-Layer Transformers
por: Musat, Tiberiu
Publicado: (2024)
por: Musat, Tiberiu
Publicado: (2024)
Pruning Weights but Not Truth: Safeguarding Truthfulness While Pruning LLMs
por: Fu, Yao, et al.
Publicado: (2025)
por: Fu, Yao, et al.
Publicado: (2025)
STUN: Structured-Then-Unstructured Pruning for Scalable MoE Pruning
por: Lee, Jaeseong, et al.
Publicado: (2024)
por: Lee, Jaeseong, et al.
Publicado: (2024)
Diversifying the Expert Knowledge for Task-Agnostic Pruning in Sparse Mixture-of-Experts
por: Zhang, Zeliang, et al.
Publicado: (2024)
por: Zhang, Zeliang, et al.
Publicado: (2024)
Beyond Linear Approximations: A Novel Pruning Approach for Attention Matrix
por: Liang, Yingyu, et al.
Publicado: (2024)
por: Liang, Yingyu, et al.
Publicado: (2024)
Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse
por: Fu, Zizhuo, et al.
Publicado: (2026)
por: Fu, Zizhuo, et al.
Publicado: (2026)
Wave-PDE Nets: Trainable Wave-Equation Layers as an Alternative to Attention
por: Vejendla, Harshil
Publicado: (2025)
por: Vejendla, Harshil
Publicado: (2025)
Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
por: Brandon, William, et al.
Publicado: (2024)
por: Brandon, William, et al.
Publicado: (2024)
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM
por: Fan, Zehao, et al.
Publicado: (2025)
por: Fan, Zehao, et al.
Publicado: (2025)
Everybody Prune Now: Structured Pruning of LLMs with only Forward Passes
por: Kolawole, Steven, et al.
Publicado: (2024)
por: Kolawole, Steven, et al.
Publicado: (2024)
From Pruning to Grafting: Dynamic Knowledge Redistribution via Learnable Layer Fusion
por: Pei, Zehua, et al.
Publicado: (2024)
por: Pei, Zehua, et al.
Publicado: (2024)
VOCABTRIM: Vocabulary Pruning for Efficient Speculative Decoding in LLMs
por: Goel, Raghavv, et al.
Publicado: (2025)
por: Goel, Raghavv, et al.
Publicado: (2025)
DELTA: Dynamic Layer-Aware Token Attention for Efficient Long-Context Reasoning
por: Zarch, Hossein Entezari, et al.
Publicado: (2025)
por: Zarch, Hossein Entezari, et al.
Publicado: (2025)
IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse
por: Bai, Yushi, et al.
Publicado: (2026)
por: Bai, Yushi, et al.
Publicado: (2026)
Rethinking Layer Redundancy: Calibration Matters More Than Search in LLM Depth Pruning
por: Kim, Minkyu, et al.
Publicado: (2026)
por: Kim, Minkyu, et al.
Publicado: (2026)
Towards Understanding the Word Sensitivity of Attention Layers: A Study via Random Features
por: Bombari, Simone, et al.
Publicado: (2024)
por: Bombari, Simone, et al.
Publicado: (2024)
Depth-Wise Attention (DWAtt): A Layer Fusion Method for Data-Efficient Classification
por: ElNokrashy, Muhammad, et al.
Publicado: (2022)
por: ElNokrashy, Muhammad, et al.
Publicado: (2022)
On Pruning State-Space LLMs
por: Ghattas, Tamer, et al.
Publicado: (2025)
por: Ghattas, Tamer, et al.
Publicado: (2025)
Causal Attention with Lookahead Keys
por: Song, Zhuoqing, et al.
Publicado: (2025)
por: Song, Zhuoqing, et al.
Publicado: (2025)
Bypass Back-propagation: Optimization-based Structural Pruning for Large Language Models via Policy Gradient
por: Gao, Yuan, et al.
Publicado: (2024)
por: Gao, Yuan, et al.
Publicado: (2024)
Ejemplares similares
-
Condense, Don't Just Prune: Enhancing Efficiency and Performance in MoE Layer Pruning
por: Cao, Mingyu, et al.
Publicado: (2024) -
Variance Control via Weight Rescaling in LLM Pre-training
por: Owen, Louis, et al.
Publicado: (2025) -
SAP: Syntactic Attention Pruning for Transformer-based Language Models
por: Lee, Tzu-Yun, et al.
Publicado: (2025) -
Twilight: Adaptive Attention Sparsity with Hierarchical Top-$p$ Pruning
por: Lin, Chaofan, et al.
Publicado: (2025) -
STAR: Spectral Truncation and Rescale for Model Merging
por: Lee, Yu-Ang, et al.
Publicado: (2025)