BESA: Pruning Large Language Models with Blockwise Parameter-Efficient Sparsity Allocation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xu, Peng, Shao, Wenqi, Chen, Mengzhao, Tang, Shitao, Zhang, Kaipeng, Gao, Peng, An, Fengwei, Qiao, Yu, Luo, Ping |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
EfficientQAT: Efficient Quantization-Aware Training for Large Language Models
von: Chen, Mengzhao, et al.
Veröffentlicht: (2024)
von: Chen, Mengzhao, et al.
Veröffentlicht: (2024)
OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models
von: Shao, Wenqi, et al.
Veröffentlicht: (2023)
von: Shao, Wenqi, et al.
Veröffentlicht: (2023)
SPP: Sparsity-Preserved Parameter-Efficient Fine-Tuning for Large Language Models
von: Lu, Xudong, et al.
Veröffentlicht: (2024)
von: Lu, Xudong, et al.
Veröffentlicht: (2024)
PrefixQuant: Eliminating Outliers by Prefixed Tokens for Large Language Models Quantization
von: Chen, Mengzhao, et al.
Veröffentlicht: (2024)
von: Chen, Mengzhao, et al.
Veröffentlicht: (2024)
DiffAgent: Fast and Accurate Text-to-Image API Selection with Large Language Model
von: Zhao, Lirui, et al.
Veröffentlicht: (2024)
von: Zhao, Lirui, et al.
Veröffentlicht: (2024)
TinyLVLM-eHub: Towards Comprehensive and Efficient Evaluation for Large Vision-Language Models
von: Shao, Wenqi, et al.
Veröffentlicht: (2023)
von: Shao, Wenqi, et al.
Veröffentlicht: (2023)
SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models
von: Liu, Dongyang, et al.
Veröffentlicht: (2024)
von: Liu, Dongyang, et al.
Veröffentlicht: (2024)
One-Shot Sensitivity-Aware Mixed Sparsity Pruning for Large Language Models
von: Shao, Hang, et al.
Veröffentlicht: (2023)
von: Shao, Hang, et al.
Veröffentlicht: (2023)
ChartAssisstant: A Universal Chart Multimodal Language Model via Chart-to-Table Pre-training and Multitask Instruction Tuning
von: Meng, Fanqing, et al.
Veröffentlicht: (2024)
von: Meng, Fanqing, et al.
Veröffentlicht: (2024)
Tree-Planner: Efficient Close-loop Task Planning with Large Language Models
von: Hu, Mengkang, et al.
Veröffentlicht: (2023)
von: Hu, Mengkang, et al.
Veröffentlicht: (2023)
MLLMs-Augmented Visual-Language Representation Learning
von: Liu, Yanqing, et al.
Veröffentlicht: (2023)
von: Liu, Yanqing, et al.
Veröffentlicht: (2023)
Sparsity Induction for Accurate Post-Training Pruning of Large Language Models
von: Jiang, Minhao, et al.
Veröffentlicht: (2026)
von: Jiang, Minhao, et al.
Veröffentlicht: (2026)
Adaptive Blockwise Search: Inference-Time Alignment for Large Language Models
von: Quamar, Mohammad Atif, et al.
Veröffentlicht: (2025)
von: Quamar, Mohammad Atif, et al.
Veröffentlicht: (2025)
Flexibly Scaling Large Language Models Contexts Through Extensible Tokenization
von: Shao, Ninglu, et al.
Veröffentlicht: (2024)
von: Shao, Ninglu, et al.
Veröffentlicht: (2024)
IG-Pruning: Input-Guided Block Pruning for Large Language Models
von: Qiao, Kangyu, et al.
Veröffentlicht: (2025)
von: Qiao, Kangyu, et al.
Veröffentlicht: (2025)
Sparsity-Accelerated Training for Large Language Models
von: Ma, Da, et al.
Veröffentlicht: (2024)
von: Ma, Da, et al.
Veröffentlicht: (2024)
Twilight: Adaptive Attention Sparsity with Hierarchical Top-$p$ Pruning
von: Lin, Chaofan, et al.
Veröffentlicht: (2025)
von: Lin, Chaofan, et al.
Veröffentlicht: (2025)
Llama2Vec: Unsupervised Adaptation of Large Language Models for Dense Retrieval
von: Liu, Zheng, et al.
Veröffentlicht: (2023)
von: Liu, Zheng, et al.
Veröffentlicht: (2023)
HiAgent: Hierarchical Working Memory Management for Solving Long-Horizon Agent Tasks with Large Language Model
von: Hu, Mengkang, et al.
Veröffentlicht: (2024)
von: Hu, Mengkang, et al.
Veröffentlicht: (2024)
BGE Landmark Embedding: A Chunking-Free Embedding Method For Retrieval Augmented Long-Context Large Language Models
von: Luo, Kun, et al.
Veröffentlicht: (2024)
von: Luo, Kun, et al.
Veröffentlicht: (2024)
NutePrune: Efficient Progressive Pruning with Numerous Teachers for Large Language Models
von: Li, Shengrui, et al.
Veröffentlicht: (2024)
von: Li, Shengrui, et al.
Veröffentlicht: (2024)
Light-PEFT: Lightening Parameter-Efficient Fine-Tuning via Early Pruning
von: Gu, Naibin, et al.
Veröffentlicht: (2024)
von: Gu, Naibin, et al.
Veröffentlicht: (2024)
Sensitivity Meets Sparsity: The Impact of Extremely Sparse Parameter Patterns on Theory-of-Mind of Large Language Models
von: Wu, Yuheng, et al.
Veröffentlicht: (2025)
von: Wu, Yuheng, et al.
Veröffentlicht: (2025)
Entropy-Based Block Pruning for Efficient Large Language Models
von: Yang, Liangwei, et al.
Veröffentlicht: (2025)
von: Yang, Liangwei, et al.
Veröffentlicht: (2025)
B-AVIBench: Towards Evaluating the Robustness of Large Vision-Language Model on Black-box Adversarial Visual-Instructions
von: Zhang, Hao, et al.
Veröffentlicht: (2024)
von: Zhang, Hao, et al.
Veröffentlicht: (2024)
TP-Eval: Tap Multimodal LLMs' Potential in Evaluation by Customizing Prompts
von: Xie, Yuxuan, et al.
Veröffentlicht: (2024)
von: Xie, Yuxuan, et al.
Veröffentlicht: (2024)
Predicting Rewards Alongside Tokens: Non-disruptive Parameter Insertion for Efficient Inference Intervention in Large Language Model
von: Yuan, Chenhan, et al.
Veröffentlicht: (2024)
von: Yuan, Chenhan, et al.
Veröffentlicht: (2024)
Learn To be Efficient: Build Structured Sparsity in Large Language Models
von: Zheng, Haizhong, et al.
Veröffentlicht: (2024)
von: Zheng, Haizhong, et al.
Veröffentlicht: (2024)
Blockwise SFT for Diffusion Language Models: Reconciling Bidirectional Attention and Autoregressive Decoding
von: Sun, Bowen, et al.
Veröffentlicht: (2025)
von: Sun, Bowen, et al.
Veröffentlicht: (2025)
Efficient Post-Training Pruning of Large Language Models with Statistical Correction
von: Yu, Peiqi, et al.
Veröffentlicht: (2026)
von: Yu, Peiqi, et al.
Veröffentlicht: (2026)
Efficient Shapley Value-based Non-Uniform Pruning of Large Language Models
von: Sun, Chuan, et al.
Veröffentlicht: (2025)
von: Sun, Chuan, et al.
Veröffentlicht: (2025)
ALPS: Attention Localization and Pruning Strategy for Efficient Alignment of Large Language Models
von: Chen, Hao, et al.
Veröffentlicht: (2025)
von: Chen, Hao, et al.
Veröffentlicht: (2025)
EvoMoE: Expert Evolution in Mixture of Experts for Multimodal Large Language Models
von: Jing, Linglin, et al.
Veröffentlicht: (2025)
von: Jing, Linglin, et al.
Veröffentlicht: (2025)
Towards Extreme Pruning of LLMs with Plug-and-Play Mixed Sparsity
von: Xu, Chi, et al.
Veröffentlicht: (2025)
von: Xu, Chi, et al.
Veröffentlicht: (2025)
One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models
von: Ye, Rongguang, et al.
Veröffentlicht: (2025)
von: Ye, Rongguang, et al.
Veröffentlicht: (2025)
Large Language Models as Foundations for Next-Gen Dense Retrieval: A Comprehensive Empirical Assessment
von: Luo, Kun, et al.
Veröffentlicht: (2024)
von: Luo, Kun, et al.
Veröffentlicht: (2024)
MINI-LLM: Memory-Efficient Structured Pruning for Large Language Models
von: Cheng, Hongrong, et al.
Veröffentlicht: (2024)
von: Cheng, Hongrong, et al.
Veröffentlicht: (2024)
KVPruner: Structural Pruning for Faster and Memory-Efficient Large Language Models
von: Lv, Bo, et al.
Veröffentlicht: (2024)
von: Lv, Bo, et al.
Veröffentlicht: (2024)
DISP-LLM: Dimension-Independent Structural Pruning for Large Language Models
von: Gao, Shangqian, et al.
Veröffentlicht: (2024)
von: Gao, Shangqian, et al.
Veröffentlicht: (2024)
Text2World: Benchmarking Large Language Models for Symbolic World Model Generation
von: Hu, Mengkang, et al.
Veröffentlicht: (2025)
von: Hu, Mengkang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
EfficientQAT: Efficient Quantization-Aware Training for Large Language Models
von: Chen, Mengzhao, et al.
Veröffentlicht: (2024) -
OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models
von: Shao, Wenqi, et al.
Veröffentlicht: (2023) -
SPP: Sparsity-Preserved Parameter-Efficient Fine-Tuning for Large Language Models
von: Lu, Xudong, et al.
Veröffentlicht: (2024) -
PrefixQuant: Eliminating Outliers by Prefixed Tokens for Large Language Models Quantization
von: Chen, Mengzhao, et al.
Veröffentlicht: (2024) -
DiffAgent: Fast and Accurate Text-to-Image API Selection with Large Language Model
von: Zhao, Lirui, et al.
Veröffentlicht: (2024)