Magnitude Pruning of Large Pretrained Transformer Models with a Mixture Gaussian Prior
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Mingxuan, Sun, Yan, Liang, Faming |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Uncertainty Quantification for Large-Scale Deep Networks via Post-StoNet Modeling
by: Sun, Yan, et al.
Published: (2025)
by: Sun, Yan, et al.
Published: (2025)
Adaptive Pruning of Pretrained Transformer via Differential Inclusions
by: Ding, Yizhuo, et al.
Published: (2025)
by: Ding, Yizhuo, et al.
Published: (2025)
Pruning and Malicious Injection: A Retraining-Free Backdoor Attack on Transformer Models
by: Zhao, Taibiao, et al.
Published: (2025)
by: Zhao, Taibiao, et al.
Published: (2025)
PDGMM-VAE: A Variational Autoencoder with Adaptive Per-Dimension Gaussian Mixture Model Priors for Nonlinear ICA
by: Wei, Yuan-Hao, et al.
Published: (2026)
by: Wei, Yuan-Hao, et al.
Published: (2026)
Structured Diffusion Models with Mixture of Gaussians as Prior Distribution
by: Jia, Nanshan, et al.
Published: (2024)
by: Jia, Nanshan, et al.
Published: (2024)
Extended Fiducial Inference: Toward an Automated Process of Statistical Inference
by: Liang, Faming, et al.
Published: (2024)
by: Liang, Faming, et al.
Published: (2024)
Federated Gaussian Mixture Models
by: Pettersson, Sophia Zhang, et al.
Published: (2025)
by: Pettersson, Sophia Zhang, et al.
Published: (2025)
Spiking Layer-Adaptive Magnitude-based Pruning
by: Wang, Junqiao, et al.
Published: (2026)
by: Wang, Junqiao, et al.
Published: (2026)
Insights into the Lottery Ticket Hypothesis and Iterative Magnitude Pruning
by: Saleem, Tausifa Jan, et al.
Published: (2024)
by: Saleem, Tausifa Jan, et al.
Published: (2024)
EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models
by: Xing, Xingrun, et al.
Published: (2025)
by: Xing, Xingrun, et al.
Published: (2025)
IPPRO: Importance-based Pruning with PRojective Offset for Magnitude-indifferent Structural Pruning
by: Jung, Jaeheun, et al.
Published: (2025)
by: Jung, Jaeheun, et al.
Published: (2025)
Causal-StoNet: Causal Inference for High-Dimensional Complex Data
by: Fang, Yaxin, et al.
Published: (2024)
by: Fang, Yaxin, et al.
Published: (2024)
Efficient Training of Large-Scale AI Models Through Federated Mixture-of-Experts: A System-Level Approach
by: Chen, Xiaobing, et al.
Published: (2025)
by: Chen, Xiaobing, et al.
Published: (2025)
Magnitude-based Neuron Pruning for Backdoor Defens
by: Li, Nan, et al.
Published: (2024)
by: Li, Nan, et al.
Published: (2024)
Combining Relevance and Magnitude for Resource-Aware DNN Pruning
by: Chiasserini, Carla Fabiana, et al.
Published: (2024)
by: Chiasserini, Carla Fabiana, et al.
Published: (2024)
Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models
by: Lu, Xudong, et al.
Published: (2024)
by: Lu, Xudong, et al.
Published: (2024)
Sparse Weight Averaging with Multiple Particles for Iterative Magnitude Pruning
by: Choi, Moonseok, et al.
Published: (2023)
by: Choi, Moonseok, et al.
Published: (2023)
End-To-End Learning of Gaussian Mixture Priors for Diffusion Sampler
by: Blessing, Denis, et al.
Published: (2025)
by: Blessing, Denis, et al.
Published: (2025)
The VampPrior Mixture Model
by: Stirn, Andrew A., et al.
Published: (2024)
by: Stirn, Andrew A., et al.
Published: (2024)
Fast Value Tracking for Deep Reinforcement Learning
by: Shih, Frank, et al.
Published: (2024)
by: Shih, Frank, et al.
Published: (2024)
Domain-Specific Pruning of Large Mixture-of-Experts Models with Few-shot Demonstrations
by: Dong, Zican, et al.
Published: (2025)
by: Dong, Zican, et al.
Published: (2025)
Certified Robustness from Approximate Gaussian Mixture Structures in Pretrained Latent Spaces
by: Emmanouilidis, Konstantinos, et al.
Published: (2026)
by: Emmanouilidis, Konstantinos, et al.
Published: (2026)
Is Complexity Required for Neural Network Pruning? A Case Study on Global Magnitude Pruning
by: Gupta, Manas, et al.
Published: (2022)
by: Gupta, Manas, et al.
Published: (2022)
Mixture of Length and Pruning Experts for Knowledge Graphs Reasoning
by: Du, Enjun, et al.
Published: (2025)
by: Du, Enjun, et al.
Published: (2025)
IDEA Prune: An Integrated Enlarge-and-Prune Pipeline in Generative Language Model Pretraining
by: Li, Yixiao, et al.
Published: (2025)
by: Li, Yixiao, et al.
Published: (2025)
Mosaic Pruning: A Hierarchical Framework for Generalizable Pruning of Mixture-of-Experts Models
by: Hu, Wentao, et al.
Published: (2025)
by: Hu, Wentao, et al.
Published: (2025)
Extended Fiducial Inference for Individual Treatment Effects via Deep Neural Networks
by: Kim, Sehwan, et al.
Published: (2025)
by: Kim, Sehwan, et al.
Published: (2025)
Efficient Reinforcement Learning with Large Language Model Priors
by: Yan, Xue, et al.
Published: (2024)
by: Yan, Xue, et al.
Published: (2024)
Uncertainty Quantification for Physics-Informed Neural Networks with Extended Fiducial Inference
by: Shih, Frank, et al.
Published: (2025)
by: Shih, Frank, et al.
Published: (2025)
FLEX-MoE: Federated Mixture-of-Experts with Load-balanced Expert Assignment for Edge Computing
by: Zhang, Boyang, et al.
Published: (2025)
by: Zhang, Boyang, et al.
Published: (2025)
Transformers as Unsupervised Learning Algorithms: A study on Gaussian Mixtures
by: Chen, Zhiheng, et al.
Published: (2025)
by: Chen, Zhiheng, et al.
Published: (2025)
On How Iterative Magnitude Pruning Discovers Local Receptive Fields in Fully Connected Neural Networks
by: Redman, William T., et al.
Published: (2024)
by: Redman, William T., et al.
Published: (2024)
Finite Neural Networks as Mixtures of Gaussian Processes: From Provable Error Bounds to Prior Selection
by: Adams, Steven, et al.
Published: (2024)
by: Adams, Steven, et al.
Published: (2024)
FedMap: Iterative Magnitude-Based Pruning for Communication-Efficient Federated Learning
by: Herzog, Alexander, et al.
Published: (2024)
by: Herzog, Alexander, et al.
Published: (2024)
Generalization Guarantees for Representation Learning via Data-Dependent Gaussian Mixture Priors
by: Sefidgaran, Milad, et al.
Published: (2025)
by: Sefidgaran, Milad, et al.
Published: (2025)
MC#: Mixture Compressor for Mixture-of-Experts Large Models
by: Huang, Wei, et al.
Published: (2025)
by: Huang, Wei, et al.
Published: (2025)
Scalable Clustering: Large Scale Unsupervised Learning of Gaussian Mixture Models with Outliers
by: Zhou, Yijia, et al.
Published: (2023)
by: Zhou, Yijia, et al.
Published: (2023)
Deep Survival Analysis for Competing Risk Modeling with Functional Covariates and Missing Data Imputation
by: Gao, Penglei, et al.
Published: (2025)
by: Gao, Penglei, et al.
Published: (2025)
Whitening Spherical Gaussian Mixtures in the Large-Dimensional Regime
by: Boudjemaa, Mohammed Racim Moussa, et al.
Published: (2025)
by: Boudjemaa, Mohammed Racim Moussa, et al.
Published: (2025)
Model Selection and Parameter Estimation of Multi-dimensional Gaussian Mixture Model
by: Liu, Xinyu, et al.
Published: (2026)
by: Liu, Xinyu, et al.
Published: (2026)
Similar Items
-
Uncertainty Quantification for Large-Scale Deep Networks via Post-StoNet Modeling
by: Sun, Yan, et al.
Published: (2025) -
Adaptive Pruning of Pretrained Transformer via Differential Inclusions
by: Ding, Yizhuo, et al.
Published: (2025) -
Pruning and Malicious Injection: A Retraining-Free Backdoor Attack on Transformer Models
by: Zhao, Taibiao, et al.
Published: (2025) -
PDGMM-VAE: A Variational Autoencoder with Adaptive Per-Dimension Gaussian Mixture Model Priors for Nonlinear ICA
by: Wei, Yuan-Hao, et al.
Published: (2026) -
Structured Diffusion Models with Mixture of Gaussians as Prior Distribution
by: Jia, Nanshan, et al.
Published: (2024)