Saved in:
| Main Authors: | Chu, Chi-Wei, Hong, Ding-Yong, Wu, Jan-Jan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2507.17301 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Performance characterisation of the 64-core SG2042 RISC-V CPU for HPC
by: Brown, Nick, et al.
Published: (2024)
by: Brown, Nick, et al.
Published: (2024)
Hybrid Dual-Batch and Cyclic Progressive Learning for Efficient Distributed Training
by: Lu, Kuan-Wei, et al.
Published: (2025)
by: Lu, Kuan-Wei, et al.
Published: (2025)
Efficient Architecture for RISC-V Vector Memory Access
by: Guan, Hongyi, et al.
Published: (2025)
by: Guan, Hongyi, et al.
Published: (2025)
Programming RISC-V accelerators via Fortran
by: Brown, Nick, et al.
Published: (2025)
by: Brown, Nick, et al.
Published: (2025)
Vitamin-V: Expanding Open-Source RISC-V Cloud Environments
by: Canal, Ramon, et al.
Published: (2024)
by: Canal, Ramon, et al.
Published: (2024)
Accelerating stencils on the Tenstorrent Grayskull RISC-V accelerator
by: Brown, Nick, et al.
Published: (2024)
by: Brown, Nick, et al.
Published: (2024)
Assessing Performance and Porting Strategies for Gravitational $N$-Body Simulations on the RISC-V-Based Tenstorrent Wormhole\textsuperscript{\texttrademark}
by: Almerol, Jenny Lynn, et al.
Published: (2026)
by: Almerol, Jenny Lynn, et al.
Published: (2026)
RISC-V for HPC: An update of where we are and main action points
by: Brown, Nick
Published: (2025)
by: Brown, Nick
Published: (2025)
Closer in the Gap: Towards Portable Performance on RISC-V Vector Processors
by: Shi, Ruimin, et al.
Published: (2026)
by: Shi, Ruimin, et al.
Published: (2026)
RISC-V for HPC: Where we are and where we need to go
by: Brown, Nick
Published: (2024)
by: Brown, Nick
Published: (2024)
Enabling an OpenStack-based cloud on top of RISC-V hardware
by: Marrón, Diego, et al.
Published: (2024)
by: Marrón, Diego, et al.
Published: (2024)
Efficient Graph Embedding at Scale: Optimizing CPU-GPU-SSD Integration
by: Li, Zhonggen, et al.
Published: (2025)
by: Li, Zhonggen, et al.
Published: (2025)
Is RISC-V ready for High Performance Computing? An evaluation of the Sophon SG2044
by: Brown, Nick
Published: (2025)
by: Brown, Nick
Published: (2025)
Fine-Grained Vectorized Merge Sorting on RISC-V: From Register to Cache
by: Zhang, Jin, et al.
Published: (2024)
by: Zhang, Jin, et al.
Published: (2024)
TURNIP: A "Nondeterministic" GPU Runtime with CPU RAM Offload
by: Ding, Zhimin, et al.
Published: (2024)
by: Ding, Zhimin, et al.
Published: (2024)
Vitamin-V: Virtual Environment and Tool-boxing for Trustworthy Development of RISC-V based Cloud Services
by: Arelakis, A., et al.
Published: (2023)
by: Arelakis, A., et al.
Published: (2023)
DiT-HC: Enabling Efficient Training of Visual Generation Model DiT on HPC-oriented CPU Cluster
by: Zhang, Jinxiao, et al.
Published: (2026)
by: Zhang, Jinxiao, et al.
Published: (2026)
Towards CXL Resilience to CPU Failures
by: Psistakis, Antonis, et al.
Published: (2026)
by: Psistakis, Antonis, et al.
Published: (2026)
Bridging Simulation and Silicon: A Study of RISC-V Hardware and FireSim Simulation
by: Barai, Atanu, et al.
Published: (2025)
by: Barai, Atanu, et al.
Published: (2025)
Monte Cimone v2: Down the Road of RISC-V High-Performance Computers
by: Venieri, Emanuele, et al.
Published: (2025)
by: Venieri, Emanuele, et al.
Published: (2025)
Monte Cimone v3: Where RISC-V Stands in High-Performance Computing
by: Venieri, Emanuele, et al.
Published: (2026)
by: Venieri, Emanuele, et al.
Published: (2026)
Preparing for HPC on RISC-V: Examining Vectorization and Distributed Performance of an Astrophyiscs Application with HPX and Kokkos
by: Diehl, Patrick, et al.
Published: (2024)
by: Diehl, Patrick, et al.
Published: (2024)
CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference
by: Li, Suyi, et al.
Published: (2024)
by: Li, Suyi, et al.
Published: (2024)
Efficient CPU-GPU Collaborative Inference for MoE-based LLMs on Memory-Limited Systems
by: Huang, En-Ming, et al.
Published: (2025)
by: Huang, En-Ming, et al.
Published: (2025)
SiPipe: Bridging the CPU-GPU Utilization Gap for Efficient Pipeline-Parallel LLM Inference
by: He, Yongchao, et al.
Published: (2025)
by: He, Yongchao, et al.
Published: (2025)
Dual-pronged deep learning preprocessing on heterogeneous platforms with CPU, Accelerator and CSD
by: Wei, Jia, et al.
Published: (2024)
by: Wei, Jia, et al.
Published: (2024)
Performance optimization of BLAS algorithms with band matrices for RISC-V processors
by: Pirova, Anna, et al.
Published: (2025)
by: Pirova, Anna, et al.
Published: (2025)
FractalSortCPU: Bandwidth-Efficient Compressed Radix Sort on CPU
by: Dang'ana, Michael
Published: (2026)
by: Dang'ana, Michael
Published: (2026)
WindVE: Collaborative CPU-NPU Vector Embedding
by: Huang, Jinqi, et al.
Published: (2025)
by: Huang, Jinqi, et al.
Published: (2025)
Combining GPU and CPU for accelerating evolutionary computing workloads
by: Eynaliyev, Rustam, et al.
Published: (2025)
by: Eynaliyev, Rustam, et al.
Published: (2025)
A Unified CPU-GPU Protocol for GNN Training
by: Lin, Yi-Chien, et al.
Published: (2024)
by: Lin, Yi-Chien, et al.
Published: (2024)
NM-SpMM: Accelerating Matrix Multiplication Using N:M Sparsity with GPGPU
by: Ma, Cong, et al.
Published: (2025)
by: Ma, Cong, et al.
Published: (2025)
MMStencil: Optimizing High-order Stencils on Multicore CPU using Matrix Unit
by: Wang, Yinuo, et al.
Published: (2025)
by: Wang, Yinuo, et al.
Published: (2025)
PruneX: A Hierarchical Communication-Efficient System for Distributed CNN Training with Structured Pruning
by: Olama, Alireza, et al.
Published: (2025)
by: Olama, Alireza, et al.
Published: (2025)
Memory-Efficient Federated Fine-Tuning of Large Language Models via Layer Pruning
by: Wu, Yebo, et al.
Published: (2025)
by: Wu, Yebo, et al.
Published: (2025)
RTop-K: Ultra-Fast Row-Wise Top-K Selection for Neural Network Acceleration on GPUs
by: Xie, Xi, et al.
Published: (2024)
by: Xie, Xi, et al.
Published: (2024)
Co-designing a Programmable RISC-V Accelerator for MPC-based Energy and Thermal Management of Many-Core HPC Processors
by: Ottaviano, Alessandro, et al.
Published: (2025)
by: Ottaviano, Alessandro, et al.
Published: (2025)
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
by: Lin, Mao, et al.
Published: (2026)
by: Lin, Mao, et al.
Published: (2026)
Justin: Hybrid CPU/Memory Elastic Scaling for Distributed Stream Processing
by: Schmitz, Donatien, et al.
Published: (2025)
by: Schmitz, Donatien, et al.
Published: (2025)
A Unified Programming Model for Heterogeneous Computing with CPU and Accelerator Technologies
by: Xiong, Yuqing
Published: (2022)
by: Xiong, Yuqing
Published: (2022)
Similar Items
-
Performance characterisation of the 64-core SG2042 RISC-V CPU for HPC
by: Brown, Nick, et al.
Published: (2024) -
Hybrid Dual-Batch and Cyclic Progressive Learning for Efficient Distributed Training
by: Lu, Kuan-Wei, et al.
Published: (2025) -
Efficient Architecture for RISC-V Vector Memory Access
by: Guan, Hongyi, et al.
Published: (2025) -
Programming RISC-V accelerators via Fortran
by: Brown, Nick, et al.
Published: (2025) -
Vitamin-V: Expanding Open-Source RISC-V Cloud Environments
by: Canal, Ramon, et al.
Published: (2024)