Efficient LLMs with AMP: Attention Heads and MLP Pruning
Fuente:
arXiv
Saved in:
| Main Authors: | Mugnaini, Leandro Giusti, Yamamoto, Bruno Lopes, de Alcantara, Lucas Lauton, Zacarias, Victor, Bollis, Edson, Pellicer, Lucas, Costa, Anna Helena Reali, Jordao, Artur |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Compressing LLMs with MoP: Mixture of Pruners
by: Yamamoto, Bruno Lopes, et al.
Published: (2026)
by: Yamamoto, Bruno Lopes, et al.
Published: (2026)
Technical Report on Text Dataset Distillation
by: Ogawa, Keith Ando, et al.
Published: (2025)
by: Ogawa, Keith Ando, et al.
Published: (2025)
Layer-wise LoRA fine-tuning: a similarity metric approach
by: Ogawa, Keith Ando, et al.
Published: (2026)
by: Ogawa, Keith Ando, et al.
Published: (2026)
Layer Pruning with Consensus: A Triple-Win Solution
by: Mugnaini, Leandro Giusti, et al.
Published: (2024)
by: Mugnaini, Leandro Giusti, et al.
Published: (2024)
Effective Layer Pruning Through Similarity Metric Perspective
by: Pons, Ian, et al.
Published: (2024)
by: Pons, Ian, et al.
Published: (2024)
The Virtues of Brevity: Avoid Overthinking in Parallel Test-Time Reasoning
by: Dinardi, Raul Cavalcante, et al.
Published: (2025)
by: Dinardi, Raul Cavalcante, et al.
Published: (2025)
Pruning Everything, Everywhere, All at Once
by: Nascimento, Gustavo Henrique do, et al.
Published: (2025)
by: Nascimento, Gustavo Henrique do, et al.
Published: (2025)
Improving Fairness in LLMs Through Testing-Time Adversaries
by: Gregio, Isabela Pereira, et al.
Published: (2025)
by: Gregio, Isabela Pereira, et al.
Published: (2025)
Weakly Supervised Attention-based Models Using Activation Maps for Citrus Mite and Insect Pest Classification
by: Bollis, Edson, et al.
Published: (2021)
by: Bollis, Edson, et al.
Published: (2021)
One Period to Rule Them All: Identifying Critical Learning Periods in Deep Networks
by: Fukase, Vinicius Yuiti, et al.
Published: (2025)
by: Fukase, Vinicius Yuiti, et al.
Published: (2025)
Realistic Market Impact Modeling for Reinforcement Learning Trading Environments
by: Abbade, Lucas Riera, et al.
Published: (2026)
by: Abbade, Lucas Riera, et al.
Published: (2026)
Developing an ESG-Oriented Large Language Model through ESG Practices
by: Assis, Gabriel, et al.
Published: (2026)
by: Assis, Gabriel, et al.
Published: (2026)
Combating the Elsagate phenomenon: Deep learning architectures for disturbing cartoons
by: Ishikawa, Akari, et al.
Published: (2019)
by: Ishikawa, Akari, et al.
Published: (2019)
EL TRAUMA ORTOPÉDICO EN LOS ANCIANOS: UNA REVISIÓN DE LA LITERATURA.
by: J. Lauton Soares
Published: (2005)
by: J. Lauton Soares
Published: (2005)
Normalização de nomes de autores em fontes de informação institucionais: proposta de um método automático de verificação de erros
by: Rogério Mugnaini
Published: (2012)
by: Rogério Mugnaini
Published: (2012)
Recuperação e impacto da produção científica na era Google: uma análise comparativa entre o Google Acadêmico e a Web of Science
by: Rogério Mugnaini
Published: (2008)
by: Rogério Mugnaini
Published: (2008)
Para un protocolo de observación etnográfica de los usos diferenciales y los modos de ver las telenovelas
by: Fabio Mugnaini
Published: (1986)
by: Fabio Mugnaini
Published: (1986)
ACESSO ABERTO E FINANCIAMENTO DA PESQUISA NO BRASIL: CARACTERÍSTICAS E TENDÊNCIAS DA PRODUÇÃO CIENTÍFICA
by: Rogério Mugnaini
Published: (2022)
by: Rogério Mugnaini
Published: (2022)
Adaptive MLP Pruning for Large Vision Transformers
by: Shen, Chengchao
Published: (2026)
by: Shen, Chengchao
Published: (2026)
Automatic Channel Pruning for Multi-Head Attention
by: Lee, Eunho, et al.
Published: (2024)
by: Lee, Eunho, et al.
Published: (2024)
Comparing Normalization Methods for Portfolio Optimization with Reinforcement Learning
by: Costa, Caio de Souza Barbosa, et al.
Published: (2025)
by: Costa, Caio de Souza Barbosa, et al.
Published: (2025)
From MLP to NeoMLP: Leveraging Self-Attention for Neural Fields
by: Kofinas, Miltiadis, et al.
Published: (2024)
by: Kofinas, Miltiadis, et al.
Published: (2024)
Extracting Small Translation Specialists from LLMs by Aggressively Pruning Experts
by: Martin, Liu O., et al.
Published: (2026)
by: Martin, Liu O., et al.
Published: (2026)
Useful woody species and its environmental availability: the case of artisanal fishermen in Itaúnas, Brazil
by: Lucas Costa Monteiro Lopes
Published: (2017)
by: Lucas Costa Monteiro Lopes
Published: (2017)
Assortment of Attention Heads: Accelerating Federated PEFT with Head Pruning and Strategic Client Selection
by: Venkatesha, Yeshwanth, et al.
Published: (2025)
by: Venkatesha, Yeshwanth, et al.
Published: (2025)
Ensembling Pruned Attention Heads For Uncertainty-Aware Efficient Transformers
by: Gabetni, Firas, et al.
Published: (2025)
by: Gabetni, Firas, et al.
Published: (2025)
The promotion and implementation of open science measures among high‐performing journals from Brazil, Mexico, Portugal, and Spain
by: Chris Fradkin, et al.
Published: (2024)
by: Chris Fradkin, et al.
Published: (2024)
SDMPrune: Self-Distillation MLP Pruning for Efficient Large Language Models
by: Zhu, Hourun, et al.
Published: (2025)
by: Zhu, Hourun, et al.
Published: (2025)
Mitigating Temporal Blindness in Kubernetes Autoscaling: An Attention-Double-LSTM Framework
by: Shaikh, Faraz, et al.
Published: (2026)
by: Shaikh, Faraz, et al.
Published: (2026)
FourierKAN outperforms MLP on Text Classification Head Fine-tuning
by: Imran, Abdullah Al, et al.
Published: (2024)
by: Imran, Abdullah Al, et al.
Published: (2024)
Data-Free Pruning of Self-Attention Layers in LLMs
by: Saikumar, Dhananjay, et al.
Published: (2025)
by: Saikumar, Dhananjay, et al.
Published: (2025)
SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning
by: Wang, Hanrui, et al.
Published: (2020)
by: Wang, Hanrui, et al.
Published: (2020)
From Random to Informed Data Selection: A Diversity-Based Approach to Optimize Human Annotation and Few-Shot Learning
by: Alcoforado, Alexandre, et al.
Published: (2024)
by: Alcoforado, Alexandre, et al.
Published: (2024)
ENTRE HISTORIA Y MEMORIA: LA PRODUCCIÓN DE LUIS A. DE HERRERA EN LOS ORÍGENES DE UN RELATO REVISIONISTA SOBRE LA GUERRA DEL PARAGUAY
by: Laura Reali
Published: (2006)
by: Laura Reali
Published: (2006)
Ridurre il dono alla donazione: Il metodo fenomenologico e la teologia secondo Jean-Luc Marion
by: Nicola Reali
Published: (2016)
by: Nicola Reali
Published: (2016)
A influência da microbiota de doadores e do armazenamento do enxerto no transplante de córnea
by: Catiusca Reali
Published: (2022)
by: Catiusca Reali
Published: (2022)
Garbage Attention in Large Language Models: BOS Sink Heads and Sink-aware Pruning
by: Sok, Jaewon, et al.
Published: (2026)
by: Sok, Jaewon, et al.
Published: (2026)
Multi-LED Classification as Pretext For Robot Heading Estimation
by: Carlotti, Nicholas, et al.
Published: (2024)
by: Carlotti, Nicholas, et al.
Published: (2024)
When Layers Play the Lottery, all Tickets Win at Initialization
by: Jordao, Artur, et al.
Published: (2023)
by: Jordao, Artur, et al.
Published: (2023)
Debiasing LLMs by Masking Unfairness-Driving Attention Heads
by: Han, Tingxu, et al.
Published: (2025)
by: Han, Tingxu, et al.
Published: (2025)
Similar Items
-
Compressing LLMs with MoP: Mixture of Pruners
by: Yamamoto, Bruno Lopes, et al.
Published: (2026) -
Technical Report on Text Dataset Distillation
by: Ogawa, Keith Ando, et al.
Published: (2025) -
Layer-wise LoRA fine-tuning: a similarity metric approach
by: Ogawa, Keith Ando, et al.
Published: (2026) -
Layer Pruning with Consensus: A Triple-Win Solution
by: Mugnaini, Leandro Giusti, et al.
Published: (2024) -
Effective Layer Pruning Through Similarity Metric Perspective
by: Pons, Ian, et al.
Published: (2024)