Efficient Column-Wise N:M Pruning on RISC-V CPU
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chu, Chi-Wei, Hong, Ding-Yong, Wu, Jan-Jan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Performance characterisation of the 64-core SG2042 RISC-V CPU for HPC
von: Brown, Nick, et al.
Veröffentlicht: (2024)
von: Brown, Nick, et al.
Veröffentlicht: (2024)
Hybrid Dual-Batch and Cyclic Progressive Learning for Efficient Distributed Training
von: Lu, Kuan-Wei, et al.
Veröffentlicht: (2025)
von: Lu, Kuan-Wei, et al.
Veröffentlicht: (2025)
Programming RISC-V accelerators via Fortran
von: Brown, Nick, et al.
Veröffentlicht: (2025)
von: Brown, Nick, et al.
Veröffentlicht: (2025)
Efficient Architecture for RISC-V Vector Memory Access
von: Guan, Hongyi, et al.
Veröffentlicht: (2025)
von: Guan, Hongyi, et al.
Veröffentlicht: (2025)
Vitamin-V: Expanding Open-Source RISC-V Cloud Environments
von: Canal, Ramon, et al.
Veröffentlicht: (2024)
von: Canal, Ramon, et al.
Veröffentlicht: (2024)
Accelerating stencils on the Tenstorrent Grayskull RISC-V accelerator
von: Brown, Nick, et al.
Veröffentlicht: (2024)
von: Brown, Nick, et al.
Veröffentlicht: (2024)
Assessing Performance and Porting Strategies for Gravitational $N$-Body Simulations on the RISC-V-Based Tenstorrent Wormhole\textsuperscript{\texttrademark}
von: Almerol, Jenny Lynn, et al.
Veröffentlicht: (2026)
von: Almerol, Jenny Lynn, et al.
Veröffentlicht: (2026)
Efficient Graph Embedding at Scale: Optimizing CPU-GPU-SSD Integration
von: Li, Zhonggen, et al.
Veröffentlicht: (2025)
von: Li, Zhonggen, et al.
Veröffentlicht: (2025)
RISC-V for HPC: An update of where we are and main action points
von: Brown, Nick
Veröffentlicht: (2025)
von: Brown, Nick
Veröffentlicht: (2025)
Closer in the Gap: Towards Portable Performance on RISC-V Vector Processors
von: Shi, Ruimin, et al.
Veröffentlicht: (2026)
von: Shi, Ruimin, et al.
Veröffentlicht: (2026)
RISC-V for HPC: Where we are and where we need to go
von: Brown, Nick
Veröffentlicht: (2024)
von: Brown, Nick
Veröffentlicht: (2024)
Enabling an OpenStack-based cloud on top of RISC-V hardware
von: Marrón, Diego, et al.
Veröffentlicht: (2024)
von: Marrón, Diego, et al.
Veröffentlicht: (2024)
Is RISC-V ready for High Performance Computing? An evaluation of the Sophon SG2044
von: Brown, Nick
Veröffentlicht: (2025)
von: Brown, Nick
Veröffentlicht: (2025)
Fine-Grained Vectorized Merge Sorting on RISC-V: From Register to Cache
von: Zhang, Jin, et al.
Veröffentlicht: (2024)
von: Zhang, Jin, et al.
Veröffentlicht: (2024)
Vitamin-V: Virtual Environment and Tool-boxing for Trustworthy Development of RISC-V based Cloud Services
von: Arelakis, A., et al.
Veröffentlicht: (2023)
von: Arelakis, A., et al.
Veröffentlicht: (2023)
DiT-HC: Enabling Efficient Training of Visual Generation Model DiT on HPC-oriented CPU Cluster
von: Zhang, Jinxiao, et al.
Veröffentlicht: (2026)
von: Zhang, Jinxiao, et al.
Veröffentlicht: (2026)
TURNIP: A "Nondeterministic" GPU Runtime with CPU RAM Offload
von: Ding, Zhimin, et al.
Veröffentlicht: (2024)
von: Ding, Zhimin, et al.
Veröffentlicht: (2024)
Bridging Simulation and Silicon: A Study of RISC-V Hardware and FireSim Simulation
von: Barai, Atanu, et al.
Veröffentlicht: (2025)
von: Barai, Atanu, et al.
Veröffentlicht: (2025)
Monte Cimone v2: Down the Road of RISC-V High-Performance Computers
von: Venieri, Emanuele, et al.
Veröffentlicht: (2025)
von: Venieri, Emanuele, et al.
Veröffentlicht: (2025)
Monte Cimone v3: Where RISC-V Stands in High-Performance Computing
von: Venieri, Emanuele, et al.
Veröffentlicht: (2026)
von: Venieri, Emanuele, et al.
Veröffentlicht: (2026)
Towards CXL Resilience to CPU Failures
von: Psistakis, Antonis, et al.
Veröffentlicht: (2026)
von: Psistakis, Antonis, et al.
Veröffentlicht: (2026)
Efficient CPU-GPU Collaborative Inference for MoE-based LLMs on Memory-Limited Systems
von: Huang, En-Ming, et al.
Veröffentlicht: (2025)
von: Huang, En-Ming, et al.
Veröffentlicht: (2025)
SiPipe: Bridging the CPU-GPU Utilization Gap for Efficient Pipeline-Parallel LLM Inference
von: He, Yongchao, et al.
Veröffentlicht: (2025)
von: He, Yongchao, et al.
Veröffentlicht: (2025)
Preparing for HPC on RISC-V: Examining Vectorization and Distributed Performance of an Astrophyiscs Application with HPX and Kokkos
von: Diehl, Patrick, et al.
Veröffentlicht: (2024)
von: Diehl, Patrick, et al.
Veröffentlicht: (2024)
CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference
von: Li, Suyi, et al.
Veröffentlicht: (2024)
von: Li, Suyi, et al.
Veröffentlicht: (2024)
Dual-pronged deep learning preprocessing on heterogeneous platforms with CPU, Accelerator and CSD
von: Wei, Jia, et al.
Veröffentlicht: (2024)
von: Wei, Jia, et al.
Veröffentlicht: (2024)
WindVE: Collaborative CPU-NPU Vector Embedding
von: Huang, Jinqi, et al.
Veröffentlicht: (2025)
von: Huang, Jinqi, et al.
Veröffentlicht: (2025)
Combining GPU and CPU for accelerating evolutionary computing workloads
von: Eynaliyev, Rustam, et al.
Veröffentlicht: (2025)
von: Eynaliyev, Rustam, et al.
Veröffentlicht: (2025)
A Unified CPU-GPU Protocol for GNN Training
von: Lin, Yi-Chien, et al.
Veröffentlicht: (2024)
von: Lin, Yi-Chien, et al.
Veröffentlicht: (2024)
MMStencil: Optimizing High-order Stencils on Multicore CPU using Matrix Unit
von: Wang, Yinuo, et al.
Veröffentlicht: (2025)
von: Wang, Yinuo, et al.
Veröffentlicht: (2025)
PruneX: A Hierarchical Communication-Efficient System for Distributed CNN Training with Structured Pruning
von: Olama, Alireza, et al.
Veröffentlicht: (2025)
von: Olama, Alireza, et al.
Veröffentlicht: (2025)
Performance optimization of BLAS algorithms with band matrices for RISC-V processors
von: Pirova, Anna, et al.
Veröffentlicht: (2025)
von: Pirova, Anna, et al.
Veröffentlicht: (2025)
Memory-Efficient Federated Fine-Tuning of Large Language Models via Layer Pruning
von: Wu, Yebo, et al.
Veröffentlicht: (2025)
von: Wu, Yebo, et al.
Veröffentlicht: (2025)
RTop-K: Ultra-Fast Row-Wise Top-K Selection for Neural Network Acceleration on GPUs
von: Xie, Xi, et al.
Veröffentlicht: (2024)
von: Xie, Xi, et al.
Veröffentlicht: (2024)
Co-designing a Programmable RISC-V Accelerator for MPC-based Energy and Thermal Management of Many-Core HPC Processors
von: Ottaviano, Alessandro, et al.
Veröffentlicht: (2025)
von: Ottaviano, Alessandro, et al.
Veröffentlicht: (2025)
Justin: Hybrid CPU/Memory Elastic Scaling for Distributed Stream Processing
von: Schmitz, Donatien, et al.
Veröffentlicht: (2025)
von: Schmitz, Donatien, et al.
Veröffentlicht: (2025)
A Unified Programming Model for Heterogeneous Computing with CPU and Accelerator Technologies
von: Xiong, Yuqing
Veröffentlicht: (2022)
von: Xiong, Yuqing
Veröffentlicht: (2022)
FractalSortCPU: Bandwidth-Efficient Compressed Radix Sort on CPU
von: Dang'ana, Michael
Veröffentlicht: (2026)
von: Dang'ana, Michael
Veröffentlicht: (2026)
Pruning Blockchain Protocols for Efficient Access Control in IoT Systems
von: Huang, Yongtao, et al.
Veröffentlicht: (2024)
von: Huang, Yongtao, et al.
Veröffentlicht: (2024)
Elastic Mixture of Rank-Wise Experts for Knowledge Reuse in Federated Fine-Tuning
von: Wu, Yebo, et al.
Veröffentlicht: (2025)
von: Wu, Yebo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Performance characterisation of the 64-core SG2042 RISC-V CPU for HPC
von: Brown, Nick, et al.
Veröffentlicht: (2024) -
Hybrid Dual-Batch and Cyclic Progressive Learning for Efficient Distributed Training
von: Lu, Kuan-Wei, et al.
Veröffentlicht: (2025) -
Programming RISC-V accelerators via Fortran
von: Brown, Nick, et al.
Veröffentlicht: (2025) -
Efficient Architecture for RISC-V Vector Memory Access
von: Guan, Hongyi, et al.
Veröffentlicht: (2025) -
Vitamin-V: Expanding Open-Source RISC-V Cloud Environments
von: Canal, Ramon, et al.
Veröffentlicht: (2024)