A Proximal Operator for Inducing 2:4-Sparsity
Fuente:
arXiv
Saved in:
| Main Authors: | Kübler, Jonas M, Wang, Yu-Xiang, Sabach, Shoham, Ansari, Navid, Kleindessner, Matthäus, Budhathoki, Kailash, Cevher, Volkan, Karypis, George |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Inference Optimization of Foundation Models on AI Accelerators
by: Park, Youngsuk, et al.
Published: (2024)
by: Park, Youngsuk, et al.
Published: (2024)
When LLMs get significantly worse: A statistical approach to detect model degradations
by: Kübler, Jonas, et al.
Published: (2026)
by: Kübler, Jonas, et al.
Published: (2026)
Block-Diagonal LoRA for Eliminating Communication Overhead in Tensor Parallel LoRA Serving
by: Wang, Xinyu, et al.
Published: (2025)
by: Wang, Xinyu, et al.
Published: (2025)
LLM-Rank: A Graph Theoretical Approach to Pruning Large Language Models
by: Hoffmann, David, et al.
Published: (2024)
by: Hoffmann, David, et al.
Published: (2024)
ProxSparse: Regularized Learning of Semi-Structured Sparsity Masks for Pretrained LLMs
by: Liu, Hongyi, et al.
Published: (2025)
by: Liu, Hongyi, et al.
Published: (2025)
Krylov Cubic Regularized Newton: A Subspace Second-Order Method with Dimension-Free Convergence Rate
by: Jiang, Ruichen, et al.
Published: (2024)
by: Jiang, Ruichen, et al.
Published: (2024)
Direct Preference Optimization with Rating Information: Practical Algorithms and Provable Gains
by: Viano, Luca, et al.
Published: (2026)
by: Viano, Luca, et al.
Published: (2026)
MADA: Meta-Adaptive Optimizers through hyper-gradient Descent
by: Ozkara, Kaan, et al.
Published: (2024)
by: Ozkara, Kaan, et al.
Published: (2024)
Revisiting SMoE Language Models by Evaluating Inefficiencies with Task Specific Expert Pruning
by: Sarkar, Soumajyoti, et al.
Published: (2024)
by: Sarkar, Soumajyoti, et al.
Published: (2024)
Directional-Clamp PPO
by: Karpel, Gilad, et al.
Published: (2025)
by: Karpel, Gilad, et al.
Published: (2025)
Displacement-Resistant Extensions of DPO with Nonconvex $f$-Divergences
by: Pipano, Idan, et al.
Published: (2026)
by: Pipano, Idan, et al.
Published: (2026)
Meaningful Causal Aggregation and Paradoxical Confounding
by: Zhu, Yuchen, et al.
Published: (2023)
by: Zhu, Yuchen, et al.
Published: (2023)
Multilinear Operator Networks
by: Cheng, Yixin, et al.
Published: (2024)
by: Cheng, Yixin, et al.
Published: (2024)
Learning with Norm Constrained, Over-parameterized, Two-layer Neural Networks
by: Liu, Fanghui, et al.
Published: (2024)
by: Liu, Fanghui, et al.
Published: (2024)
SAMPa: Sharpness-aware Minimization Parallelized
by: Xie, Wanyun, et al.
Published: (2024)
by: Xie, Wanyun, et al.
Published: (2024)
MaD-Mix: Multi-Modal Data Mixtures via Latent Space Coupling for Vision-Language Model Training
by: Xie, Wanyun, et al.
Published: (2026)
by: Xie, Wanyun, et al.
Published: (2026)
IL-SOAR : Imitation Learning with Soft Optimistic Actor cRitic
by: Viel, Stefano, et al.
Published: (2025)
by: Viel, Stefano, et al.
Published: (2025)
Imitation Learning in Discounted Linear MDPs without exploration assumptions
by: Viano, Luca, et al.
Published: (2024)
by: Viano, Luca, et al.
Published: (2024)
Learning the Target Network in Function Space
by: Asadi, Kavosh, et al.
Published: (2024)
by: Asadi, Kavosh, et al.
Published: (2024)
Chameleon: A Flexible Data-mixing Framework for Language Model Pretraining and Finetuning
by: Xie, Wanyun, et al.
Published: (2025)
by: Xie, Wanyun, et al.
Published: (2025)
Efficient Large Language Model Inference with Neural Block Linearization
by: Erdogan, Mete, et al.
Published: (2025)
by: Erdogan, Mete, et al.
Published: (2025)
Stable Nonconvex-Nonconcave Training via Linear Interpolation
by: Pethick, Thomas, et al.
Published: (2023)
by: Pethick, Thomas, et al.
Published: (2023)
Extending Input Contexts of Language Models through Training on Segmented Sequences
by: Karypis, Petros, et al.
Published: (2023)
by: Karypis, Petros, et al.
Published: (2023)
μP$^2$: Effective Sharpness Aware Minimization Requires Layerwise Perturbation Scaling
by: Haas, Moritz, et al.
Published: (2024)
by: Haas, Moritz, et al.
Published: (2024)
FisherSFT: Data-Efficient Supervised Fine-Tuning of Language Models Using Information Gain
by: Deb, Rohan, et al.
Published: (2025)
by: Deb, Rohan, et al.
Published: (2025)
MT-NAM: An Efficient and Adaptive Model for Epileptic Seizure Detection
by: Afzal, Arshia, et al.
Published: (2025)
by: Afzal, Arshia, et al.
Published: (2025)
Split the Differences, Pool the Rest: Provably Efficient Multi-Objective Imitation
by: Sheebaelhamd, Ziyad, et al.
Published: (2026)
by: Sheebaelhamd, Ziyad, et al.
Published: (2026)
Optimistic Dual Averaging Unifies Modern Optimizers
by: Pethick, Thomas, et al.
Published: (2026)
by: Pethick, Thomas, et al.
Published: (2026)
Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution
by: Barla, Adam, et al.
Published: (2026)
by: Barla, Adam, et al.
Published: (2026)
High-Dimensional Kernel Methods under Covariate Shift: Data-Dependent Implicit Regularization
by: Chen, Yihang, et al.
Published: (2024)
by: Chen, Yihang, et al.
Published: (2024)
Easy Data Unlearning Bench
by: Rinberg, Roy, et al.
Published: (2026)
by: Rinberg, Roy, et al.
Published: (2026)
Adversarial Training for Defense Against Label Poisoning Attacks
by: Bal, Melis Ilayda, et al.
Published: (2025)
by: Bal, Melis Ilayda, et al.
Published: (2025)
ESLM: Risk-Averse Selective Language Modeling for Efficient Pretraining
by: Bal, Melis Ilayda, et al.
Published: (2025)
by: Bal, Melis Ilayda, et al.
Published: (2025)
Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe
by: Thekumparampil, Kiran Koshy, et al.
Published: (2024)
by: Thekumparampil, Kiran Koshy, et al.
Published: (2024)
Adversarial Training Should Be Cast as a Non-Zero-Sum Game
by: Robey, Alexander, et al.
Published: (2023)
by: Robey, Alexander, et al.
Published: (2023)
Matching Multiple Experts: On the Exploitability of Multi-Agent Imitation Learning
by: Bergerault, Antoine, et al.
Published: (2026)
by: Bergerault, Antoine, et al.
Published: (2026)
DualKV: Shared-Prompt Flash Attention for Efficient RL Training with Large Rollouts and Long Contexts
by: Gai, Jiading, et al.
Published: (2026)
by: Gai, Jiading, et al.
Published: (2026)
Proximal-IMH: Proximal Posterior Proposals for Independent Metropolis-Hastings with Approximate Operators
by: Chen, Youguang, et al.
Published: (2026)
by: Chen, Youguang, et al.
Published: (2026)
GNN-RAG: Graph Neural Retrieval for Large Language Model Reasoning
by: Mavromatis, Costas, et al.
Published: (2024)
by: Mavromatis, Costas, et al.
Published: (2024)
Learning Equilibria from Data: Provably Efficient Multi-Agent Imitation Learning
by: Freihaut, Till, et al.
Published: (2025)
by: Freihaut, Till, et al.
Published: (2025)
Similar Items
-
Inference Optimization of Foundation Models on AI Accelerators
by: Park, Youngsuk, et al.
Published: (2024) -
When LLMs get significantly worse: A statistical approach to detect model degradations
by: Kübler, Jonas, et al.
Published: (2026) -
Block-Diagonal LoRA for Eliminating Communication Overhead in Tensor Parallel LoRA Serving
by: Wang, Xinyu, et al.
Published: (2025) -
LLM-Rank: A Graph Theoretical Approach to Pruning Large Language Models
by: Hoffmann, David, et al.
Published: (2024) -
ProxSparse: Regularized Learning of Semi-Structured Sparsity Masks for Pretrained LLMs
by: Liu, Hongyi, et al.
Published: (2025)