Accelerating HDC-CNN Hybrid Models Using Custom Instructions on RISC-V GPUs
Fuente:
arXiv
Saved in:
| Main Authors: | Matsumi, Wakuto, Mian, Riaz-Ul-Haque |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GPU Volume Rendering with Hierarchical Compression Using VDB
by: Zellmann, Stefan, et al.
Published: (2025)
by: Zellmann, Stefan, et al.
Published: (2025)
Towards Real-Time Neural Volumetric Rendering on Mobile Devices: A Measurement Study
by: Wang, Zhe, et al.
Published: (2024)
by: Wang, Zhe, et al.
Published: (2024)
AIvaluateXR: An Evaluation Framework for on-Device AI in XR with Benchmarking Results
by: Khan, Dawar, et al.
Published: (2025)
by: Khan, Dawar, et al.
Published: (2025)
Capsule: Efficient Player Isolation for Datacenters
by: Du, Zhouheng, et al.
Published: (2025)
by: Du, Zhouheng, et al.
Published: (2025)
GALE: Leveraging Heterogeneous Systems for Efficient Unstructured Mesh Data Analysis
by: Liu, Guoxi, et al.
Published: (2025)
by: Liu, Guoxi, et al.
Published: (2025)
An efficient GPU approach for designing 3D cultural heritage information systems
by: López, Luis, et al.
Published: (2025)
by: López, Luis, et al.
Published: (2025)
Fast Sparse Matrix Permutation for Mesh-Based Direct Solvers
by: Zarebavami, Behrooz, et al.
Published: (2026)
by: Zarebavami, Behrooz, et al.
Published: (2026)
RAFI -- A Ray/Work Forwarding Infrastructure for Data Parallel Multi-Node/Multi-GPU Computing
by: Wald, Ingo, et al.
Published: (2026)
by: Wald, Ingo, et al.
Published: (2026)
MSz: An Efficient Parallel Algorithm for Correcting Morse-Smale Segmentations in Error-Bounded Lossy Compressors
by: Li, Yuxiao, et al.
Published: (2024)
by: Li, Yuxiao, et al.
Published: (2024)
Barrier-Augmented Lagrangian for GPU-based Elastodynamic Contact
by: Guo, Dewen, et al.
Published: (2024)
by: Guo, Dewen, et al.
Published: (2024)
Exploiting ray tracing technology through OptiX to compute particle interactions with cutoff in a 3D environment on GPU
by: David, Algis, et al.
Published: (2024)
by: David, Algis, et al.
Published: (2024)
NeRF-XL: Scaling NeRFs with Multiple GPUs
by: Li, Ruilong, et al.
Published: (2024)
by: Li, Ruilong, et al.
Published: (2024)
Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUs
by: Cui, Shengkun, et al.
Published: (2025)
by: Cui, Shengkun, et al.
Published: (2025)
Accelerating Large Language Model Training with Hybrid GPU-based Compression
by: Xu, Lang, et al.
Published: (2024)
by: Xu, Lang, et al.
Published: (2024)
Opara: Exploiting Operator Parallelism for Expediting DNN Inference on GPUs
by: Chen, Aodong, et al.
Published: (2023)
by: Chen, Aodong, et al.
Published: (2023)
Training LLMs with Fault Tolerant HSDP on 100,000 GPUs
by: Salpekar, Omkar, et al.
Published: (2026)
by: Salpekar, Omkar, et al.
Published: (2026)
A Nonlinear Hash-based Optimization Method for SpMV on GPUs
by: Yan, Chen, et al.
Published: (2025)
by: Yan, Chen, et al.
Published: (2025)
Accelerating stencils on the Tenstorrent Grayskull RISC-V accelerator
by: Brown, Nick, et al.
Published: (2024)
by: Brown, Nick, et al.
Published: (2024)
Characterizing Performance-Energy Trade-offs of Large Language Models in Multi-Request Workflows
by: Ifath, Md. Monzurul Amin, et al.
Published: (2026)
by: Ifath, Md. Monzurul Amin, et al.
Published: (2026)
Astra: Efficient and Money-saving Automatic Parallel Strategies Search on Heterogeneous GPUs
by: Wang, Peiran, et al.
Published: (2025)
by: Wang, Peiran, et al.
Published: (2025)
Arkade: k-Nearest Neighbor Search With Non-Euclidean Distances using GPU Ray Tracing
by: Mandarapu, Durga, et al.
Published: (2023)
by: Mandarapu, Durga, et al.
Published: (2023)
A Few GPUs, A Whole Lotta Scale: Faithful LLM Training Emulation with PrismLLM
by: Xi, Shaoke, et al.
Published: (2026)
by: Xi, Shaoke, et al.
Published: (2026)
Lightning Grasp: High Performance Procedural Grasp Synthesis with Contact Fields
by: Yin, Zhao-Heng, et al.
Published: (2025)
by: Yin, Zhao-Heng, et al.
Published: (2025)
Paris: A Decentralized Trained Open-Weight Diffusion Model
by: Jiang, Zhiying, et al.
Published: (2025)
by: Jiang, Zhiying, et al.
Published: (2025)
Energy Efficient Federated Learning with Hyperdimensional Computing (HDC)
by: Ding, Yahao, et al.
Published: (2026)
by: Ding, Yahao, et al.
Published: (2026)
DiffPhD: A Unified Differentiable Solver for Projective Heterogeneous Materials in Elastodynamics with Contact-Rich GPU-Acceleration
by: Lai, Shih-Yu, et al.
Published: (2026)
by: Lai, Shih-Yu, et al.
Published: (2026)
Confidential Computing on NVIDIA Hopper GPUs: A Performance Benchmark Study
by: Zhu, Jianwei, et al.
Published: (2024)
by: Zhu, Jianwei, et al.
Published: (2024)
FlashMLA-ETAP: Efficient Transpose Attention Pipeline for Accelerating MLA Inference on NVIDIA H20 GPUs
by: Dege, Pengcuo, et al.
Published: (2025)
by: Dege, Pengcuo, et al.
Published: (2025)
LR-CNN: Lightweight Row-centric Convolutional Neural Network Training for Memory Reduction
by: Wang, Zhigang, et al.
Published: (2024)
by: Wang, Zhigang, et al.
Published: (2024)
Understand and Accelerate Memory Processing Pipeline for Large Language Model Inference
by: He, Zifan, et al.
Published: (2026)
by: He, Zifan, et al.
Published: (2026)
Accelerating Maximal Biclique Enumeration on GPUs
by: Hsieh, Chou-Ying, et al.
Published: (2024)
by: Hsieh, Chou-Ying, et al.
Published: (2024)
SpecEE: Accelerating Large Language Model Inference with Speculative Early Exiting
by: Xu, Jiaming, et al.
Published: (2025)
by: Xu, Jiaming, et al.
Published: (2025)
A Survey on Large Language Model Acceleration based on KV Cache Management
by: Li, Haoyang, et al.
Published: (2024)
by: Li, Haoyang, et al.
Published: (2024)
iOS as Acceleration
by: Chen, Alexander K.
Published: (2025)
by: Chen, Alexander K.
Published: (2025)
Prompt-Aware Scheduling for Efficient Text-to-Image Inferencing System
by: Agarwal, Shubham, et al.
Published: (2025)
by: Agarwal, Shubham, et al.
Published: (2025)
Distributed Simulation of Large Multi-body Systems
by: Kale, Manas, et al.
Published: (2024)
by: Kale, Manas, et al.
Published: (2024)
HPC-based Solvers of Minimisation Problems for Signal Processing
by: Cammarasana, Simone, et al.
Published: (2023)
by: Cammarasana, Simone, et al.
Published: (2023)
A 4D Hybrid Algorithm to Scale Parallel Training to Thousands of GPUs
by: Singh, Siddharth, et al.
Published: (2023)
by: Singh, Siddharth, et al.
Published: (2023)
OrchMLLM: Orchestrate Multimodal Data with Batch Post-Balancing to Accelerate Multimodal Large Language Model Training
by: Zheng, Yijie, et al.
Published: (2025)
by: Zheng, Yijie, et al.
Published: (2025)
Accelerating LLM Inference with Precomputed Query Storage
by: Park, Jay H., et al.
Published: (2025)
by: Park, Jay H., et al.
Published: (2025)
Similar Items
-
GPU Volume Rendering with Hierarchical Compression Using VDB
by: Zellmann, Stefan, et al.
Published: (2025) -
Towards Real-Time Neural Volumetric Rendering on Mobile Devices: A Measurement Study
by: Wang, Zhe, et al.
Published: (2024) -
AIvaluateXR: An Evaluation Framework for on-Device AI in XR with Benchmarking Results
by: Khan, Dawar, et al.
Published: (2025) -
Capsule: Efficient Player Isolation for Datacenters
by: Du, Zhouheng, et al.
Published: (2025) -
GALE: Leveraging Heterogeneous Systems for Efficient Unstructured Mesh Data Analysis
by: Liu, Guoxi, et al.
Published: (2025)