Motivating Next-Gen Accelerators with Flexible (N:M) Activation Sparsity via Benchmarking Lightweight Post-Training Sparsification Approaches
Fuente:
arXiv
Saved in:
| Main Authors: | Alanova, Shirin, Kazistova, Kristina, Galaeva, Ekaterina, Kostromina, Alina, Smirnov, Vladimir, Dmitry, Redko, Dontsov, Alexey, Zhelnin, Maxim, Burnaev, Evgeny, Shvetsov, Egor |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GIFT-SW: Gaussian noise Injected Fine-Tuning of Salient Weights for LLMs
by: Zhelnin, Maxim, et al.
Published: (2024)
by: Zhelnin, Maxim, et al.
Published: (2024)
How to model Human Actions distribution with Event Sequence Data
by: Surkov, Egor, et al.
Published: (2025)
by: Surkov, Egor, et al.
Published: (2025)
MLEM: Generative and Contrastive Learning as Distinct Modalities for Event Sequences
by: Moskvoretskii, Viktor, et al.
Published: (2024)
by: Moskvoretskii, Viktor, et al.
Published: (2024)
Faster and Memory-Efficient Training of Sequential Recommendation Models for Large Catalogs
by: Zhelnin, Maxim, et al.
Published: (2025)
by: Zhelnin, Maxim, et al.
Published: (2025)
EBES: Easy Benchmarking for Event Sequences
by: Osin, Dmitry, et al.
Published: (2024)
by: Osin, Dmitry, et al.
Published: (2024)
A model for proppant dynamics in a perforated wellbore
by: Dontsov, Egor
Published: (2023)
by: Dontsov, Egor
Published: (2023)
Analysis of a constant height hydraulic fracture
by: Dontsov, Egor
Published: (2021)
by: Dontsov, Egor
Published: (2021)
QuantNAS for super resolution: searching for efficient quantization-friendly architectures against quantization noise
by: Shvetsov, Egor, et al.
Published: (2022)
by: Shvetsov, Egor, et al.
Published: (2022)
From 2:4 to 8:16 sparsity patterns in LLMs for Outliers and Weights with Variance Correction
by: Maximov, Egor, et al.
Published: (2025)
by: Maximov, Egor, et al.
Published: (2025)
Investigating the Impact of Quantization Methods on the Safety and Reliability of Large Language Models
by: Kharinaev, Artyom, et al.
Published: (2025)
by: Kharinaev, Artyom, et al.
Published: (2025)
From Internal Representations to Text Quality: A Geometric Approach to LLM Evaluation
by: Yusupov, Viacheslav, et al.
Published: (2025)
by: Yusupov, Viacheslav, et al.
Published: (2025)
Cross-Lingual Jailbreak Detection via Semantic Codebooks
by: Alanova, Shirin, et al.
Published: (2026)
by: Alanova, Shirin, et al.
Published: (2026)
StageOpt technical write-up
by: Dontsov, Egor, et al.
Published: (2023)
by: Dontsov, Egor, et al.
Published: (2023)
Beyond Early-Token Bias: Model-Specific and Language-Specific Position Effects in Multilingual LLMs
by: Menschikov, Mikhail, et al.
Published: (2025)
by: Menschikov, Mikhail, et al.
Published: (2025)
Lascoux polynomials and subdivisions of Gelfand-Zetlin polytopes
by: Presnova, Ekaterina, et al.
Published: (2023)
by: Presnova, Ekaterina, et al.
Published: (2023)
SeqNAS: Neural Architecture Search for Event Sequence Classification
by: Udovichenko, Igor, et al.
Published: (2024)
by: Udovichenko, Igor, et al.
Published: (2024)
Spatial Re-parameterization for N:M Sparsity
by: Zhang, Yuxin, et al.
Published: (2023)
by: Zhang, Yuxin, et al.
Published: (2023)
NM-SpMM: Accelerating Matrix Multiplication Using N:M Sparsity with GPGPU
by: Ma, Cong, et al.
Published: (2025)
by: Ma, Cong, et al.
Published: (2025)
Accelerating Newton-Schulz Iteration for Orthogonalization via Chebyshev-type Polynomials
by: Grishina, Ekaterina, et al.
Published: (2025)
by: Grishina, Ekaterina, et al.
Published: (2025)
Minimum Variance Unbiased N:M Sparsity for the Neural Gradients
by: Chmiel, Brian, et al.
Published: (2022)
by: Chmiel, Brian, et al.
Published: (2022)
Toward Efficient Permutation for Hierarchical N:M Sparsity on GPUs
by: Yu, Seungmin, et al.
Published: (2024)
by: Yu, Seungmin, et al.
Published: (2024)
Progressive Gradient Flow for Robust N:M Sparsity Training in Transformers
by: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Published: (2024)
by: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Published: (2024)
MaxQ: Multi-Axis Query for N:M Sparsity Network
by: Xiang, Jingyang, et al.
Published: (2023)
by: Xiang, Jingyang, et al.
Published: (2023)
Toward Lightweight and Fast Decoders for Diffusion Models in Image and Video Generation
by: Buzovkin, Alexey, et al.
Published: (2025)
by: Buzovkin, Alexey, et al.
Published: (2025)
ELSA: Exploiting Layer-wise N:M Sparsity for Vision Transformer Acceleration
by: Huang, Ning-Chi, et al.
Published: (2024)
by: Huang, Ning-Chi, et al.
Published: (2024)
An upscaling based three parameter elastic anisotropy model
by: Dontsov, E. V.
Published: (2023)
by: Dontsov, E. V.
Published: (2023)
A continuous fracture front tracking algorithm with multi layer tip elements (MuLTipEl) for a plane strain hydraulic fracture
by: Dontsov, E. V.
Published: (2021)
by: Dontsov, E. V.
Published: (2021)
The machine learning platform for developers of large systems
by: Naikov, Alexey, et al.
Published: (2025)
by: Naikov, Alexey, et al.
Published: (2025)
Partial order on involutive permutations and double Schubert cells
by: Smirnov, Evgeny
Published: (2024)
by: Smirnov, Evgeny
Published: (2024)
Friezes and continued fractions
by: Smirnov, Evgeny
Published: (2025)
by: Smirnov, Evgeny
Published: (2025)
MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs
by: Sun, Yan, et al.
Published: (2025)
by: Sun, Yan, et al.
Published: (2025)
Designing Social Learning
by: Smirnov, Aleksei, et al.
Published: (2024)
by: Smirnov, Aleksei, et al.
Published: (2024)
Multi-Agentic Approach for History Matching of Oil Reservoirs
by: Samigullin, Linar, et al.
Published: (2026)
by: Samigullin, Linar, et al.
Published: (2026)
Light Schrödinger Bridge
by: Korotin, Alexander, et al.
Published: (2023)
by: Korotin, Alexander, et al.
Published: (2023)
The Density of Cross-Persistence Diagrams and Its Applications
by: Mironenko, Alexander, et al.
Published: (2026)
by: Mironenko, Alexander, et al.
Published: (2026)
Amber Pruner: Leveraging N:M Activation Sparsity for Efficient Prefill in Large Language Models
by: An, Tai, et al.
Published: (2025)
by: An, Tai, et al.
Published: (2025)
Quantized topological transport mediated by the long-range couplings
by: Lebedeva, Ekaterina S., et al.
Published: (2025)
by: Lebedeva, Ekaterina S., et al.
Published: (2025)
Q-RAG: Long Context Multi-step Retrieval via Value-based Embedder Training
by: Sorokin, Artyom, et al.
Published: (2025)
by: Sorokin, Artyom, et al.
Published: (2025)
Towards Foundation Time Series Model: To Synthesize Or Not To Synthesize?
by: Kuvshinova, Kseniia, et al.
Published: (2024)
by: Kuvshinova, Kseniia, et al.
Published: (2024)
Tsururu: A Python-based Time Series Forecasting Strategies Library
by: Kostromina, Alina, et al.
Published: (2025)
by: Kostromina, Alina, et al.
Published: (2025)
Similar Items
-
GIFT-SW: Gaussian noise Injected Fine-Tuning of Salient Weights for LLMs
by: Zhelnin, Maxim, et al.
Published: (2024) -
How to model Human Actions distribution with Event Sequence Data
by: Surkov, Egor, et al.
Published: (2025) -
MLEM: Generative and Contrastive Learning as Distinct Modalities for Event Sequences
by: Moskvoretskii, Viktor, et al.
Published: (2024) -
Faster and Memory-Efficient Training of Sequential Recommendation Models for Large Catalogs
by: Zhelnin, Maxim, et al.
Published: (2025) -
EBES: Easy Benchmarking for Event Sequences
by: Osin, Dmitry, et al.
Published: (2024)