TSENOR: Highly-Efficient Algorithm for Finding Transposable N:M Sparse Masks
Fuente:
arXiv
Guardado en:
| Autores principales: | Meng, Xiang, Makni, Mehdi, Mazumder, Rahul |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
3BASiL: An Algorithmic Framework for Sparse plus Low-Rank Compression of LLMs
por: Makni, Mehdi, et al.
Publicado: (2026)
por: Makni, Mehdi, et al.
Publicado: (2026)
Toward Efficient Permutation for Hierarchical N:M Sparsity on GPUs
por: Yu, Seungmin, et al.
Publicado: (2024)
por: Yu, Seungmin, et al.
Publicado: (2024)
ADMM-Q: An Improved Hessian-based Weight Quantizer for Post-Training Quantization of Large Language Models
por: Lucas, Ryan, et al.
Publicado: (2026)
por: Lucas, Ryan, et al.
Publicado: (2026)
PermLLM: Learnable Channel Permutation for N:M Sparse Large Language Models
por: Zou, Lancheng, et al.
Publicado: (2025)
por: Zou, Lancheng, et al.
Publicado: (2025)
HASSLE-free: A unified Framework for Sparse plus Low-Rank Matrix Decomposition for LLMs
por: Makni, Mehdi, et al.
Publicado: (2025)
por: Makni, Mehdi, et al.
Publicado: (2025)
Minimum Variance Unbiased N:M Sparsity for the Neural Gradients
por: Chmiel, Brian, et al.
Publicado: (2022)
por: Chmiel, Brian, et al.
Publicado: (2022)
Amber Pruner: Leveraging N:M Activation Sparsity for Efficient Prefill in Large Language Models
por: An, Tai, et al.
Publicado: (2025)
por: An, Tai, et al.
Publicado: (2025)
E-Sparse: Boosting the Large Language Model Inference through Entropy-based N:M Sparsity
por: Li, Yun, et al.
Publicado: (2023)
por: Li, Yun, et al.
Publicado: (2023)
An Optimization Framework for Differentially Private Sparse Fine-Tuning
por: Makni, Mehdi, et al.
Publicado: (2025)
por: Makni, Mehdi, et al.
Publicado: (2025)
Unified Kernel-Segregated Transpose Convolution Operation
por: Tida, Vijay Srinivas, et al.
Publicado: (2025)
por: Tida, Vijay Srinivas, et al.
Publicado: (2025)
Motivating Next-Gen Accelerators with Flexible (N:M) Activation Sparsity via Benchmarking Lightweight Post-Training Sparsification Approaches
por: Alanova, Shirin, et al.
Publicado: (2025)
por: Alanova, Shirin, et al.
Publicado: (2025)
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator
por: Ramachandran, Akshat, et al.
Publicado: (2025)
por: Ramachandran, Akshat, et al.
Publicado: (2025)
Random Masking Finds Winning Tickets for Parameter Efficient Fine-tuning
por: Xu, Jing, et al.
Publicado: (2024)
por: Xu, Jing, et al.
Publicado: (2024)
ALPS: Improved Optimization for Highly Sparse One-Shot Pruning for Large Language Models
por: Meng, Xiang, et al.
Publicado: (2024)
por: Meng, Xiang, et al.
Publicado: (2024)
Neuro-symbolic Action Masking for Deep Reinforcement Learning
por: Han, Shuai, et al.
Publicado: (2026)
por: Han, Shuai, et al.
Publicado: (2026)
i-Mask: An Intelligent Mask for Breath-Driven Activity Recognition
por: Sinha, Ashutosh Kumar, et al.
Publicado: (2025)
por: Sinha, Ashutosh Kumar, et al.
Publicado: (2025)
Improving Robustness In Sparse Autoencoders via Masked Regularization
por: Narayanaswamy, Vivek, et al.
Publicado: (2026)
por: Narayanaswamy, Vivek, et al.
Publicado: (2026)
SAFE: Finding Sparse and Flat Minima to Improve Pruning
por: Lee, Dongyeop, et al.
Publicado: (2025)
por: Lee, Dongyeop, et al.
Publicado: (2025)
Sparse Autoencoders Do Not Find Canonical Units of Analysis
por: Leask, Patrick, et al.
Publicado: (2025)
por: Leask, Patrick, et al.
Publicado: (2025)
SparseSwaps: Tractable LLM Pruning Mask Refinement at Scale
por: Zimmer, Max, et al.
Publicado: (2025)
por: Zimmer, Max, et al.
Publicado: (2025)
An Approach to Variable Clustering: K-means in Transposed Data and its Relationship with Principal Component Analysis
por: Saquicela, Victor, et al.
Publicado: (2025)
por: Saquicela, Victor, et al.
Publicado: (2025)
Finding Clustering Algorithms in the Transformer Architecture
por: Clarkson, Kenneth L., et al.
Publicado: (2025)
por: Clarkson, Kenneth L., et al.
Publicado: (2025)
Trainable Dynamic Mask Sparse Attention
por: Shi, Jingze, et al.
Publicado: (2025)
por: Shi, Jingze, et al.
Publicado: (2025)
Progressive Sparse Attention: Algorithm and System Co-design for Efficient Attention in LLM Serving
por: Zhou, Qihui, et al.
Publicado: (2025)
por: Zhou, Qihui, et al.
Publicado: (2025)
SPADE: Faster Drug Discovery by Learning from Sparse Data
por: Nandakumar, Rahul, et al.
Publicado: (2026)
por: Nandakumar, Rahul, et al.
Publicado: (2026)
Robust Batch-Level Query Routing for Large Language Models under Cost and Capacity Constraints
por: Markovic-Voronov, Jelena, et al.
Publicado: (2026)
por: Markovic-Voronov, Jelena, et al.
Publicado: (2026)
Finding Belief Geometries with Sparse Autoencoders
por: Levinson, Matthew
Publicado: (2026)
por: Levinson, Matthew
Publicado: (2026)
SparseDM: Toward Sparse Efficient Diffusion Models
por: Wang, Kafeng, et al.
Publicado: (2024)
por: Wang, Kafeng, et al.
Publicado: (2024)
SparsePO: Controlling Preference Alignment of LLMs via Sparse Token Masks
por: Christopoulou, Fenia, et al.
Publicado: (2024)
por: Christopoulou, Fenia, et al.
Publicado: (2024)
SDMixer: Sparse Dual-Mixer for Time Series Forecasting
por: Ao, Xiang
Publicado: (2026)
por: Ao, Xiang
Publicado: (2026)
Scaling Algorithm Distillation for Continuous Control with Mamba
por: Beaussant, Samuel, et al.
Publicado: (2025)
por: Beaussant, Samuel, et al.
Publicado: (2025)
HISA: Efficient Hierarchical Indexing for Fine-Grained Sparse Attention
por: Xu, Yufei, et al.
Publicado: (2026)
por: Xu, Yufei, et al.
Publicado: (2026)
Equivariant Masked Position Prediction for Efficient Molecular Representation
por: An, Junyi, et al.
Publicado: (2025)
por: An, Junyi, et al.
Publicado: (2025)
One-Pass to Reason: Token Duplication and Block-Sparse Mask for Efficient Fine-Tuning on Multi-Turn Reasoning
por: Goru, Ritesh, et al.
Publicado: (2025)
por: Goru, Ritesh, et al.
Publicado: (2025)
FPGA Co-Design for Efficient N:M Sparse and Quantized Model Inference
por: Hsieh, Fen-Yu, et al.
Publicado: (2025)
por: Hsieh, Fen-Yu, et al.
Publicado: (2025)
Sparse High Rank Adapters
por: Bhardwaj, Kartikeya, et al.
Publicado: (2024)
por: Bhardwaj, Kartikeya, et al.
Publicado: (2024)
ACPO: A Policy Optimization Algorithm for Average MDPs with Constraints
por: Agnihotri, Akhil, et al.
Publicado: (2023)
por: Agnihotri, Akhil, et al.
Publicado: (2023)
Efficient On-Device Agents via Adaptive Context Management
por: Vijayvargiya, Sanidhya, et al.
Publicado: (2025)
por: Vijayvargiya, Sanidhya, et al.
Publicado: (2025)
FlashMLA-ETAP: Efficient Transpose Attention Pipeline for Accelerating MLA Inference on NVIDIA H20 GPUs
por: Dege, Pengcuo, et al.
Publicado: (2025)
por: Dege, Pengcuo, et al.
Publicado: (2025)
Topology-Aware Revival for Efficient Sparse Training
por: Jin, Meiling, et al.
Publicado: (2026)
por: Jin, Meiling, et al.
Publicado: (2026)
Ejemplares similares
-
3BASiL: An Algorithmic Framework for Sparse plus Low-Rank Compression of LLMs
por: Makni, Mehdi, et al.
Publicado: (2026) -
Toward Efficient Permutation for Hierarchical N:M Sparsity on GPUs
por: Yu, Seungmin, et al.
Publicado: (2024) -
ADMM-Q: An Improved Hessian-based Weight Quantizer for Post-Training Quantization of Large Language Models
por: Lucas, Ryan, et al.
Publicado: (2026) -
PermLLM: Learnable Channel Permutation for N:M Sparse Large Language Models
por: Zou, Lancheng, et al.
Publicado: (2025) -
HASSLE-free: A unified Framework for Sparse plus Low-Rank Matrix Decomposition for LLMs
por: Makni, Mehdi, et al.
Publicado: (2025)