Progressive Gradient Flow for Robust N:M Sparsity Training in Transformers
Fuente:
arXiv
Salvato in:
| Autori principali: | Bambhaniya, Abhimanyu Rajeshkumar, Yazdanbakhsh, Amir, Subramanian, Suvinay, Kao, Sheng-Chun, Agrawal, Shivani, Evci, Utku, Krishna, Tushar |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MIST: A Co-Design Framework for Heterogeneous, Multi-Stage LLM Inference
di: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Pubblicazione: (2025)
di: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Pubblicazione: (2025)
Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models
di: Bambhaniya, Abhimanyu, et al.
Pubblicazione: (2024)
di: Bambhaniya, Abhimanyu, et al.
Pubblicazione: (2024)
Enabling Unstructured Sparse Acceleration on Structured Sparse Accelerators
di: Jeong, Geonhwa, et al.
Pubblicazione: (2024)
di: Jeong, Geonhwa, et al.
Pubblicazione: (2024)
Beyond Moore's Law: Harnessing the Redshift of Generative AI with Effective Hardware-Software Co-Design
di: Yazdanbakhsh, Amir
Pubblicazione: (2025)
di: Yazdanbakhsh, Amir
Pubblicazione: (2025)
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator
di: Ramachandran, Akshat, et al.
Pubblicazione: (2025)
di: Ramachandran, Akshat, et al.
Pubblicazione: (2025)
Characterizing VLA Models: Identifying the Action Generation Bottleneck for Edge AI Architectures
di: Vishwanathan, Manoj, et al.
Pubblicazione: (2026)
di: Vishwanathan, Manoj, et al.
Pubblicazione: (2026)
FG-Attn: Leveraging Fine-Grained Sparsity In Diffusion Transformers
di: Durvasula, Sankeerth, et al.
Pubblicazione: (2025)
di: Durvasula, Sankeerth, et al.
Pubblicazione: (2025)
Scaling Multi-Node Mixture-of-Experts Inference Using Expert Activation Patterns
di: Bambhaniya, Abhimanyu, et al.
Pubblicazione: (2026)
di: Bambhaniya, Abhimanyu, et al.
Pubblicazione: (2026)
Tao: Re-Thinking DL-based Microarchitecture Simulation
di: Pandey, Santosh, et al.
Pubblicazione: (2024)
di: Pandey, Santosh, et al.
Pubblicazione: (2024)
FEATHER: A Reconfigurable Accelerator with Data Reordering Support for Low-Cost On-Chip Dataflow Switching
di: Tong, Jianming, et al.
Pubblicazione: (2024)
di: Tong, Jianming, et al.
Pubblicazione: (2024)
MCMComm: Hardware-Software Co-Optimization for End-to-End Communication in Multi-Chip-Modules
di: Raj, Ritik, et al.
Pubblicazione: (2025)
di: Raj, Ritik, et al.
Pubblicazione: (2025)
FRED: Flexible REduction-Distribution Interconnect and Communication Implementation for Wafer-Scale Distributed Training of DNN Models
di: Rashidi, Saeed, et al.
Pubblicazione: (2024)
di: Rashidi, Saeed, et al.
Pubblicazione: (2024)
Understanding the Performance Horizon of the Latest ML Workloads with NonGEMM Workloads
di: Karami, Rachid, et al.
Pubblicazione: (2024)
di: Karami, Rachid, et al.
Pubblicazione: (2024)
PipeOrgan: Efficient Inter-operation Pipelining with Flexible Spatial Organization and Interconnects
di: Garg, Raveesh, et al.
Pubblicazione: (2024)
di: Garg, Raveesh, et al.
Pubblicazione: (2024)
Axon: A novel systolic array architecture for improved run time and energy efficient GeMM and Conv operation with on-chip im2col
di: Nayan, Md Mizanur Rahaman, et al.
Pubblicazione: (2025)
di: Nayan, Md Mizanur Rahaman, et al.
Pubblicazione: (2025)
SCALE-Sim TPU: Validating and Extending SCALE-Sim for TPUs
di: Dang, Jingtian, et al.
Pubblicazione: (2026)
di: Dang, Jingtian, et al.
Pubblicazione: (2026)
CogSys: Efficient and Scalable Neurosymbolic Cognition System via Algorithm-Hardware Co-Design
di: Wan, Zishen, et al.
Pubblicazione: (2025)
di: Wan, Zishen, et al.
Pubblicazione: (2025)
SATA: Sparsity-Aware Scheduling for Selective Token Attention
di: Fan, Zhenkun, et al.
Pubblicazione: (2026)
di: Fan, Zhenkun, et al.
Pubblicazione: (2026)
In-Storage Domain-Specific Acceleration for Serverless Computing
di: Mahapatra, Rohan, et al.
Pubblicazione: (2023)
di: Mahapatra, Rohan, et al.
Pubblicazione: (2023)
HASS: Hardware-Aware Sparsity Search for Dataflow DNN Accelerator
di: Yu, Zhewen, et al.
Pubblicazione: (2024)
di: Yu, Zhewen, et al.
Pubblicazione: (2024)
Multilayer Dataflow: Orchestrate Butterfly Sparsity to Accelerate Attention Computation
di: Wu, Haibin, et al.
Pubblicazione: (2024)
di: Wu, Haibin, et al.
Pubblicazione: (2024)
Prosperity: Accelerating Spiking Neural Networks via Product Sparsity
di: Wei, Chiyue, et al.
Pubblicazione: (2025)
di: Wei, Chiyue, et al.
Pubblicazione: (2025)
FPGA Co-Design for Efficient N:M Sparse and Quantized Model Inference
di: Hsieh, Fen-Yu, et al.
Pubblicazione: (2025)
di: Hsieh, Fen-Yu, et al.
Pubblicazione: (2025)
PULSE: Parametric Hardware Units for Low-power Sparsity-Aware Convolution Engine
di: Aliyev, Ilkin, et al.
Pubblicazione: (2024)
di: Aliyev, Ilkin, et al.
Pubblicazione: (2024)
VIKIN: A Reconfigurable Accelerator for KANs and MLPs with Two-Stage Sparsity Support
di: Ou, Wenhui, et al.
Pubblicazione: (2026)
di: Ou, Wenhui, et al.
Pubblicazione: (2026)
Kratos: An FPGA Benchmark for Unrolled DNNs with Fine-Grained Sparsity and Mixed Precision
di: Dai, Xilai, et al.
Pubblicazione: (2024)
di: Dai, Xilai, et al.
Pubblicazione: (2024)
OneDSE: A Unified Microprocessor Metric Prediction and Design Space Exploration Framework
di: Raj, Ritik, et al.
Pubblicazione: (2025)
di: Raj, Ritik, et al.
Pubblicazione: (2025)
MINISA: Minimal Instruction Set Architecture for Next-gen Reconfigurable Inference Accelerator
di: Tong, Jianming, et al.
Pubblicazione: (2026)
di: Tong, Jianming, et al.
Pubblicazione: (2026)
Sparsity-Aware Streaming SNN Accelerator with Output-Channel Dataflow for Automatic Modulation Classification
di: Yang, Kuilian, et al.
Pubblicazione: (2026)
di: Yang, Kuilian, et al.
Pubblicazione: (2026)
Exploring the Sparsity-Quantization Interplay on a Novel Hybrid SNN Event-Driven Architecture
di: Aliyev, Ilkin, et al.
Pubblicazione: (2024)
di: Aliyev, Ilkin, et al.
Pubblicazione: (2024)
LogicSparse: Enabling Engine-Free Unstructured Sparsity for Quantised Deep-learning Accelerators
di: Li, Changhong, et al.
Pubblicazione: (2025)
di: Li, Changhong, et al.
Pubblicazione: (2025)
Commercial Evaluation of Zero-Skipping MAC Design for Bit Sparsity Exploitation in DL Inference
di: Nair, Harideep, et al.
Pubblicazione: (2024)
di: Nair, Harideep, et al.
Pubblicazione: (2024)
Towards Efficient SRAM-PIM Architecture Design by Exploiting Unstructured Bit-Level Sparsity
di: Duan, Cenlin, et al.
Pubblicazione: (2024)
di: Duan, Cenlin, et al.
Pubblicazione: (2024)
TENET: An Efficient Sparsity-Aware LUT-Centric Architecture for Ternary LLM Inference On Edge
di: Huang, Zhirui, et al.
Pubblicazione: (2025)
di: Huang, Zhirui, et al.
Pubblicazione: (2025)
DeMM: A Decoupled Matrix Multiplication Engine Supporting Relaxed Structured Sparsity
di: Peltekis, Christodoulos, et al.
Pubblicazione: (2024)
di: Peltekis, Christodoulos, et al.
Pubblicazione: (2024)
SCALE-Sim v3: A modular cycle-accurate systolic accelerator simulator for end-to-end system analysis
di: Raj, Ritik, et al.
Pubblicazione: (2025)
di: Raj, Ritik, et al.
Pubblicazione: (2025)
RNM-TD3: N:M Semi-structured Sparse Reinforcement Learning From Scratch
di: Vrce, Isam, et al.
Pubblicazione: (2026)
di: Vrce, Isam, et al.
Pubblicazione: (2026)
Cross-Layer Design of Vector-Symbolic Computing: Bridging Cognition and Brain-Inspired Hardware Acceleration
di: Du, Shuting, et al.
Pubblicazione: (2025)
di: Du, Shuting, et al.
Pubblicazione: (2025)
Efficient SRAM-PIM Co-design by Joint Exploration of Value-Level and Bit-Level Sparsity
di: Duan, Cenlin, et al.
Pubblicazione: (2025)
di: Duan, Cenlin, et al.
Pubblicazione: (2025)
SOFA: A Compute-Memory Optimized Sparsity Accelerator via Cross-Stage Coordinated Tiling
di: Wang, Huizheng, et al.
Pubblicazione: (2024)
di: Wang, Huizheng, et al.
Pubblicazione: (2024)
Documenti analoghi
-
MIST: A Co-Design Framework for Heterogeneous, Multi-Stage LLM Inference
di: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Pubblicazione: (2025) -
Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models
di: Bambhaniya, Abhimanyu, et al.
Pubblicazione: (2024) -
Enabling Unstructured Sparse Acceleration on Structured Sparse Accelerators
di: Jeong, Geonhwa, et al.
Pubblicazione: (2024) -
Beyond Moore's Law: Harnessing the Redshift of Generative AI with Effective Hardware-Software Co-Design
di: Yazdanbakhsh, Amir
Pubblicazione: (2025) -
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator
di: Ramachandran, Akshat, et al.
Pubblicazione: (2025)