PermLLM: Learnable Channel Permutation for N:M Sparse Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zou, Lancheng, Yin, Shuo, Pei, Zehua, Ho, Tsung-Yi, Farnia, Farzan, Yu, Bei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PermLLM: Private Inference of Large Language Models within 3 Seconds under WAN
by: Zheng, Fei, et al.
Published: (2024)
by: Zheng, Fei, et al.
Published: (2024)
CPPL: A Circuit Prompt Programming Language
by: Yin, Shuo, et al.
Published: (2026)
by: Yin, Shuo, et al.
Published: (2026)
Sparse Domain Transfer via Elastic Net Regularization
by: Zhang, Jingwei, et al.
Published: (2024)
by: Zhang, Jingwei, et al.
Published: (2024)
Toward Efficient Permutation for Hierarchical N:M Sparsity on GPUs
by: Yu, Seungmin, et al.
Published: (2024)
by: Yu, Seungmin, et al.
Published: (2024)
MoreauPruner: Robust Pruning of Large Language Models against Weight Perturbations
by: Wang, Zixiao, et al.
Published: (2024)
by: Wang, Zixiao, et al.
Published: (2024)
PipeRTL: Timing-Aware Pipeline Optimization at IR-Level for RTL Generation
by: Yin, Shuo, et al.
Published: (2026)
by: Yin, Shuo, et al.
Published: (2026)
APOLLO: Automated LLM and Lean Collaboration for Advanced Formal Reasoning
by: Ospanov, Azim, et al.
Published: (2025)
by: Ospanov, Azim, et al.
Published: (2025)
The Maximum von Neumann Entropy Principle: Theory and Applications in Machine Learning
by: Wu, Youqi, et al.
Published: (2026)
by: Wu, Youqi, et al.
Published: (2026)
DAK-UCB: Diversity-Aware Prompt Routing for LLMs and Generative Models
by: Jafari, Donya, et al.
Published: (2026)
by: Jafari, Donya, et al.
Published: (2026)
Do Vendi Scores Converge with Finite Samples? Truncated Vendi Score for Finite-Sample Convergence Guarantees
by: Ospanov, Azim, et al.
Published: (2024)
by: Ospanov, Azim, et al.
Published: (2024)
Gaussian Smoothing in Saliency Maps: The Stability-Fidelity Trade-Off in Neural Network Interpretability
by: Ye, Zhuorui, et al.
Published: (2024)
by: Ye, Zhuorui, et al.
Published: (2024)
Certified Adversarial Robustness via Partition-based Randomized Smoothing
by: Goli, Hossein, et al.
Published: (2024)
by: Goli, Hossein, et al.
Published: (2024)
Analytical FFN-to-MoE Restructuring via Activation Pattern Analysis
by: Pei, Zehua, et al.
Published: (2025)
by: Pei, Zehua, et al.
Published: (2025)
On the Fragility of AI-Based Channel Decoders under Small Channel Perturbations
by: Lei, Haoyu, et al.
Published: (2026)
by: Lei, Haoyu, et al.
Published: (2026)
PAK-UCB Contextual Bandit: An Online Learning Approach to Prompt-Aware Selection of Generative Models and LLMs
by: Hu, Xiaoyan, et al.
Published: (2024)
by: Hu, Xiaoyan, et al.
Published: (2024)
A Multi-Armed Bandit Approach to Online Selection and Evaluation of Generative Models
by: Hu, Xiaoyan, et al.
Published: (2024)
by: Hu, Xiaoyan, et al.
Published: (2024)
An Information Theoretic Approach to Interaction-Grounded Learning
by: Hu, Xiaoyan, et al.
Published: (2024)
by: Hu, Xiaoyan, et al.
Published: (2024)
On the Distributed Evaluation of Generative Models
by: Wang, Zixiao, et al.
Published: (2023)
by: Wang, Zixiao, et al.
Published: (2023)
Sparse by Rule: Probability-Based N:M Pruning for Spiking Neural Networks
by: Ye, Shuhan, et al.
Published: (2025)
by: Ye, Shuhan, et al.
Published: (2025)
E-Sparse: Boosting the Large Language Model Inference through Entropy-based N:M Sparsity
by: Li, Yun, et al.
Published: (2023)
by: Li, Yun, et al.
Published: (2023)
When Exploration Comes for Free with Mixture-Greedy: Do we need UCB in Diversity-Aware Multi-Armed Bandits?
by: Nia, Bahar Dibaei, et al.
Published: (2026)
by: Nia, Bahar Dibaei, et al.
Published: (2026)
Interpretation of Neural Networks is Susceptible to Universal Adversarial Perturbations
by: Oskouie, Haniyeh Ehsani, et al.
Published: (2022)
by: Oskouie, Haniyeh Ehsani, et al.
Published: (2022)
ChatPattern: Layout Pattern Customization via Natural Language
by: Wang, Zixiao, et al.
Published: (2024)
by: Wang, Zixiao, et al.
Published: (2024)
MixPE: Quantization and Hardware Co-design for Efficient LLM Inference
by: Zhang, Yu, et al.
Published: (2024)
by: Zhang, Yu, et al.
Published: (2024)
FPGA Co-Design for Efficient N:M Sparse and Quantized Model Inference
by: Hsieh, Fen-Yu, et al.
Published: (2025)
by: Hsieh, Fen-Yu, et al.
Published: (2025)
From Pruning to Grafting: Dynamic Knowledge Redistribution via Learnable Layer Fusion
by: Pei, Zehua, et al.
Published: (2024)
by: Pei, Zehua, et al.
Published: (2024)
PromptWise: Online Learning for Cost-Aware Prompt Assignment in Generative Models
by: Hu, Xiaoyan, et al.
Published: (2025)
by: Hu, Xiaoyan, et al.
Published: (2025)
Exposing Diversity Bias in Deep Generative Models: Statistical Origins and Correction of Diversity Error
by: Farnia, Farzan, et al.
Published: (2026)
by: Farnia, Farzan, et al.
Published: (2026)
When Kernels Multiply, Clusters Unify: Fusing Embeddings with the Kronecker Product
by: Wu, Youqi, et al.
Published: (2025)
by: Wu, Youqi, et al.
Published: (2025)
Structured Gradient-based Interpretations via Norm-Regularized Adversarial Training
by: Gong, Shizhan, et al.
Published: (2024)
by: Gong, Shizhan, et al.
Published: (2024)
Stability and Generalization in Free Adversarial Training
by: Cheng, Xiwei, et al.
Published: (2024)
by: Cheng, Xiwei, et al.
Published: (2024)
Scendi Score: Prompt-Aware Diversity Evaluation via Schur Complement of CLIP Embeddings
by: Ospanov, Azim, et al.
Published: (2024)
by: Ospanov, Azim, et al.
Published: (2024)
miniF2F-Lean Revisited: Reviewing Limitations and Charting a Path Forward
by: Ospanov, Azim, et al.
Published: (2025)
by: Ospanov, Azim, et al.
Published: (2025)
PromptSplit: Revealing Prompt-Level Disagreement in Generative Models
by: Lotfian, Mehdi, et al.
Published: (2026)
by: Lotfian, Mehdi, et al.
Published: (2026)
On the Hardness of Sampling from Mixture Distributions via Langevin Dynamics
by: Cheng, Xiwei, et al.
Published: (2024)
by: Cheng, Xiwei, et al.
Published: (2024)
On the Inductive Biases of Demographic Parity-based Fair Learning Algorithms
by: Lei, Haoyu, et al.
Published: (2024)
by: Lei, Haoyu, et al.
Published: (2024)
LoopPerm-CPD: A Robust Loop Permutation Framework for Automatic Multiple Change-Point Detection in Longitudinal Data
by: Sun, Xuejun, et al.
Published: (2026)
by: Sun, Xuejun, et al.
Published: (2026)
TSENOR: Highly-Efficient Algorithm for Finding Transposable N:M Sparse Masks
by: Meng, Xiang, et al.
Published: (2025)
by: Meng, Xiang, et al.
Published: (2025)
NM-SpMM: Accelerating Matrix Multiplication Using N:M Sparsity with GPGPU
by: Ma, Cong, et al.
Published: (2025)
by: Ma, Cong, et al.
Published: (2025)
NeuPerm: Disrupting Malware Hidden in Neural Network Parameters by Leveraging Permutation Symmetry
by: Gilkarov, Daniel, et al.
Published: (2025)
by: Gilkarov, Daniel, et al.
Published: (2025)
Similar Items
-
PermLLM: Private Inference of Large Language Models within 3 Seconds under WAN
by: Zheng, Fei, et al.
Published: (2024) -
CPPL: A Circuit Prompt Programming Language
by: Yin, Shuo, et al.
Published: (2026) -
Sparse Domain Transfer via Elastic Net Regularization
by: Zhang, Jingwei, et al.
Published: (2024) -
Toward Efficient Permutation for Hierarchical N:M Sparsity on GPUs
by: Yu, Seungmin, et al.
Published: (2024) -
MoreauPruner: Robust Pruning of Large Language Models against Weight Perturbations
by: Wang, Zixiao, et al.
Published: (2024)