CRISP: Hybrid Structured Sparsity for Class-aware Model Pruning
Fuente:
arXiv
Saved in:
| Main Authors: | Aggarwal, Shivam, Binici, Kuluhan, Mitra, Tulika |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Shedding the Bits: Pushing the Boundaries of Quantization with Minifloats on FPGAs
by: Aggarwal, Shivam, et al.
Published: (2023)
by: Aggarwal, Shivam, et al.
Published: (2023)
Mix-and-Match Pruning: Globally Guided Layer-Wise Sparsification of DNNs
by: Monachan, Danial, et al.
Published: (2026)
by: Monachan, Danial, et al.
Published: (2026)
SQ-DM: Accelerating Diffusion Models with Aggressive Quantization and Temporal Sparsity
by: Fan, Zichen, et al.
Published: (2025)
by: Fan, Zichen, et al.
Published: (2025)
Neural Architecture Search of Hybrid Models for NPU-CIM Heterogeneous AR/VR Devices
by: Zhao, Yiwei, et al.
Published: (2024)
by: Zhao, Yiwei, et al.
Published: (2024)
QOC: Quantum On-Chip Training with Parameter Shift and Gradient Pruning
by: Wang, Hanrui, et al.
Published: (2022)
by: Wang, Hanrui, et al.
Published: (2022)
Ditto: Accelerating Diffusion Model via Temporal Value Similarity
by: Kim, Sungbin, et al.
Published: (2025)
by: Kim, Sungbin, et al.
Published: (2025)
Benchmarking Deep Learning Models on NVIDIA Jetson Nano for Real-Time Systems: An Empirical Investigation
by: Swaminathan, Tushar Prasanna, et al.
Published: (2024)
by: Swaminathan, Tushar Prasanna, et al.
Published: (2024)
AHCQ-SAM: Toward Accurate and Hardware-Compatible Post-Training Segment Anything Model Quantization
by: Zhang, Wenlun, et al.
Published: (2025)
by: Zhang, Wenlun, et al.
Published: (2025)
ViM-Q: Scalable Algorithm-Hardware Co-Design for Vision Mamba Model Inference on FPGA
by: Lyu, Shengzhe, et al.
Published: (2026)
by: Lyu, Shengzhe, et al.
Published: (2026)
FG-Attn: Leveraging Fine-Grained Sparsity In Diffusion Transformers
by: Durvasula, Sankeerth, et al.
Published: (2025)
by: Durvasula, Sankeerth, et al.
Published: (2025)
Stella Nera: A Differentiable Maddness-Based Hardware Accelerator for Efficient Approximate Matrix Multiplication
by: Schönleber, Jannis, et al.
Published: (2023)
by: Schönleber, Jannis, et al.
Published: (2023)
NeuralFuse: Learning to Recover the Accuracy of Access-Limited Neural Network Inference in Low-Voltage Regimes
by: Sun, Hao-Lun, et al.
Published: (2023)
by: Sun, Hao-Lun, et al.
Published: (2023)
HARFLOW3D: A Latency-Oriented 3D-CNN Accelerator Toolflow for HAR on FPGA Devices
by: Toupas, Petros, et al.
Published: (2023)
by: Toupas, Petros, et al.
Published: (2023)
Energy Efficient Exact and Approximate Systolic Array Architecture for Matrix Multiplication
by: Jaswal, Pragun, et al.
Published: (2025)
by: Jaswal, Pragun, et al.
Published: (2025)
Smaller, Faster, Cheaper: Architectural Designs for Efficient Machine Learning
by: Walton, Steven
Published: (2025)
by: Walton, Steven
Published: (2025)
SMOF: Streaming Modern CNNs on FPGAs with Smart Off-Chip Eviction
by: Toupas, Petros, et al.
Published: (2024)
by: Toupas, Petros, et al.
Published: (2024)
Performance Analysis of Edge and In-Sensor AI Processors: A Comparative Review
by: Capogrosso, Luigi, et al.
Published: (2026)
by: Capogrosso, Luigi, et al.
Published: (2026)
QUILL: An Algorithm-Architecture Co-Design for Cache-Local Deformable Attention
by: Oh, Hyunwoo, et al.
Published: (2025)
by: Oh, Hyunwoo, et al.
Published: (2025)
Neuro-Channel Networks: A Multiplication-Free Architecture by Biological Signal Transmission
by: Mete, Emrah, et al.
Published: (2026)
by: Mete, Emrah, et al.
Published: (2026)
ViTCoD: Vision Transformer Acceleration via Dedicated Algorithm and Accelerator Co-Design
by: You, Haoran, et al.
Published: (2022)
by: You, Haoran, et al.
Published: (2022)
End-to-end fully-binarized network design: from Generic Learned Thermometer to Block Pruning
by: Nguyen, Thien, et al.
Published: (2025)
by: Nguyen, Thien, et al.
Published: (2025)
Fine-Tuning Small Language Models for Domain-Specific AI: An Edge AI Perspective
by: Aralimatti, Rakshit, et al.
Published: (2025)
by: Aralimatti, Rakshit, et al.
Published: (2025)
Accelerating 3D Gaussian Splatting with Neural Sorting and Axis-Oriented Rasterization
by: Wang, Zhican, et al.
Published: (2025)
by: Wang, Zhican, et al.
Published: (2025)
RaGNNarok: A Light-Weight Graph Neural Network for Enhancing Radar Point Clouds on Unmanned Ground Vehicles
by: Hunt, David, et al.
Published: (2025)
by: Hunt, David, et al.
Published: (2025)
On Latency Predictors for Neural Architecture Search
by: Akhauri, Yash, et al.
Published: (2024)
by: Akhauri, Yash, et al.
Published: (2024)
HALO: Hardware-aware quantization with low critical-path-delay weights for LLM acceleration
by: Juneja, Rohan, et al.
Published: (2025)
by: Juneja, Rohan, et al.
Published: (2025)
FabGPT: An Efficient Large Multimodal Model for Complex Wafer Defect Knowledge Queries
by: Jiang, Yuqi, et al.
Published: (2024)
by: Jiang, Yuqi, et al.
Published: (2024)
fpgaHART: A toolflow for throughput-oriented acceleration of 3D CNNs for HAR onto FPGAs
by: Toupas, Petros, et al.
Published: (2023)
by: Toupas, Petros, et al.
Published: (2023)
Rapid-INR: Storage Efficient CPU-free DNN Training Using Implicit Neural Representation
by: Chen, Hanqiu, et al.
Published: (2023)
by: Chen, Hanqiu, et al.
Published: (2023)
FMM-X3D: FPGA-based modeling and mapping of X3D for Human Action Recognition
by: Toupas, Petros, et al.
Published: (2023)
by: Toupas, Petros, et al.
Published: (2023)
Real-World Deployment of a Lane Change Prediction Architecture Based on Knowledge Graph Embeddings and Bayesian Inference
by: Manzour, M., et al.
Published: (2025)
by: Manzour, M., et al.
Published: (2025)
TSLA: A Task-Specific Learning Adaptation for Semantic Segmentation on Autonomous Vehicles Platform
by: Liu, Jun, et al.
Published: (2025)
by: Liu, Jun, et al.
Published: (2025)
SageAttention2++: A More Efficient Implementation of SageAttention2
by: Zhang, Jintao, et al.
Published: (2025)
by: Zhang, Jintao, et al.
Published: (2025)
Real-Time Semantic Segmentation of Aerial Images Using an Embedded U-Net: A Comparison of CPU, GPU, and FPGA Workflows
by: Posso, Julien, et al.
Published: (2025)
by: Posso, Julien, et al.
Published: (2025)
Design Insights and Comparative Evaluation of a Hardware-Based Cooperative Perception Architecture for Lane Change Prediction
by: Manzour, Mohamed, et al.
Published: (2025)
by: Manzour, Mohamed, et al.
Published: (2025)
Model Quantization and Hardware Acceleration for Vision Transformers: A Comprehensive Survey
by: Du, Dayou, et al.
Published: (2024)
by: Du, Dayou, et al.
Published: (2024)
TerEffic: Highly Efficient Ternary LLM Inference on FPGA
by: Yin, Chenyang, et al.
Published: (2025)
by: Yin, Chenyang, et al.
Published: (2025)
Towards Efficient Deployment of Hybrid SNNs on Neuromorphic and Edge AI Hardware
by: Seekings, James, et al.
Published: (2024)
by: Seekings, James, et al.
Published: (2024)
Vision Transformers on the Edge: A Comprehensive Survey of Model Compression and Acceleration Strategies
by: Saha, Shaibal, et al.
Published: (2025)
by: Saha, Shaibal, et al.
Published: (2025)
SF-MMCN: Low-Power Sever Flow Multi-Mode Diffusion Model Accelerator
by: Hsu, Huan-Ke, et al.
Published: (2024)
by: Hsu, Huan-Ke, et al.
Published: (2024)
Similar Items
-
Shedding the Bits: Pushing the Boundaries of Quantization with Minifloats on FPGAs
by: Aggarwal, Shivam, et al.
Published: (2023) -
Mix-and-Match Pruning: Globally Guided Layer-Wise Sparsification of DNNs
by: Monachan, Danial, et al.
Published: (2026) -
SQ-DM: Accelerating Diffusion Models with Aggressive Quantization and Temporal Sparsity
by: Fan, Zichen, et al.
Published: (2025) -
Neural Architecture Search of Hybrid Models for NPU-CIM Heterogeneous AR/VR Devices
by: Zhao, Yiwei, et al.
Published: (2024) -
QOC: Quantum On-Chip Training with Parameter Shift and Gradient Pruning
by: Wang, Hanrui, et al.
Published: (2022)