Minuet: Accelerating 3D Sparse Convolutions on GPUs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Jiacheng, Giannoula, Christina, Wu, Jun, Elhoushi, Mostafa, Gleeson, James, Pekhimenko, Gennady |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DCC: Data-Centric Compilation of Machine Learning Kernels for Processing-In-Memory Architectures
von: Yang, Peiming, et al.
Veröffentlicht: (2025)
von: Yang, Peiming, et al.
Veröffentlicht: (2025)
SwiftFusion: Scalable Sequence Parallelism for Distributed Inference of Diffusion Transformers on GPUs
von: Yang, Jiacheng, et al.
Veröffentlicht: (2026)
von: Yang, Jiacheng, et al.
Veröffentlicht: (2026)
Spira: Exploiting Voxel Data Structural Properties for Efficient Sparse Convolution in Point Cloud Networks
von: Adamopoulos, Dionysios, et al.
Veröffentlicht: (2025)
von: Adamopoulos, Dionysios, et al.
Veröffentlicht: (2025)
PyGim: An Efficient Graph Neural Network Library for Real Processing-In-Memory Architectures
von: Giannoula, Christina, et al.
Veröffentlicht: (2024)
von: Giannoula, Christina, et al.
Veröffentlicht: (2024)
How to Rent GPUs on a Budget
von: Li, Zhouzi, et al.
Veröffentlicht: (2024)
von: Li, Zhouzi, et al.
Veröffentlicht: (2024)
Fusing Depthwise and Pointwise Convolutions for Efficient Inference on GPUs
von: Qararyah, Fareed, et al.
Veröffentlicht: (2024)
von: Qararyah, Fareed, et al.
Veröffentlicht: (2024)
GROMACS Unplugged: How Power Capping and Frequency Shapes Performance on GPUs
von: Afzal, Ayesha, et al.
Veröffentlicht: (2025)
von: Afzal, Ayesha, et al.
Veröffentlicht: (2025)
Preliminary report: Initial evaluation of StdPar implementations on AMD GPUs for HPC
von: Lin, Wei-Chen, et al.
Veröffentlicht: (2024)
von: Lin, Wei-Chen, et al.
Veröffentlicht: (2024)
A Comprehensive Analysis of Process Energy Consumption on Multi-Socket Systems with GPUs
von: León-Vega, Luis G., et al.
Veröffentlicht: (2024)
von: León-Vega, Luis G., et al.
Veröffentlicht: (2024)
Cyclic Data Streaming on GPUs for Short Range Stencils Applied to Molecular Dynamics
von: Rose, Martin, et al.
Veröffentlicht: (2025)
von: Rose, Martin, et al.
Veröffentlicht: (2025)
Bringing Auto-tuning to HIP: Analysis of Tuning Impact and Difficulty on AMD and Nvidia GPUs
von: Lurati, Milo, et al.
Veröffentlicht: (2024)
von: Lurati, Milo, et al.
Veröffentlicht: (2024)
Alya towards Exascale: Optimal OpenACC Performance of the Navier-Stokes Finite Element Assembly on GPUs
von: Owen, Herbert, et al.
Veröffentlicht: (2024)
von: Owen, Herbert, et al.
Veröffentlicht: (2024)
Is Sparse Matrix Reordering Effective for Sparse Matrix-Vector Multiplication?
von: Asudeh, Omid, et al.
Veröffentlicht: (2025)
von: Asudeh, Omid, et al.
Veröffentlicht: (2025)
Kubernetes in Action: Exploring the Performance of Kubernetes Distributions in the Cloud
von: Aqasizade, Hossein, et al.
Veröffentlicht: (2024)
von: Aqasizade, Hossein, et al.
Veröffentlicht: (2024)
SHIRO: Near-Optimal Communication Strategies for Distributed Sparse Matrix Multiplication
von: Zhuang, Chen, et al.
Veröffentlicht: (2025)
von: Zhuang, Chen, et al.
Veröffentlicht: (2025)
Staging Blocked Evaluation over Structured Sparse Matrices
von: Das, Pratyush, et al.
Veröffentlicht: (2024)
von: Das, Pratyush, et al.
Veröffentlicht: (2024)
PASTA: A Modular Program Analysis Tool Framework for Accelerators
von: Lin, Mao, et al.
Veröffentlicht: (2026)
von: Lin, Mao, et al.
Veröffentlicht: (2026)
LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind
von: Zhang, Li, et al.
Veröffentlicht: (2025)
von: Zhang, Li, et al.
Veröffentlicht: (2025)
Architecture Specific Generation of Large Scale Lattice Boltzmann Methods for Sparse Complex Geometries
von: Suffa, Philipp, et al.
Veröffentlicht: (2024)
von: Suffa, Philipp, et al.
Veröffentlicht: (2024)
An Auto-tuning Method for Run-time Data Transformation for Sparse Matrix-Vector Multiplication
von: Katagiri, Takahiro, et al.
Veröffentlicht: (2024)
von: Katagiri, Takahiro, et al.
Veröffentlicht: (2024)
Accelerating Gaussian beam tracing method with dynamic parallelism on graphics processing units
von: Sheng, Zhang, et al.
Veröffentlicht: (2025)
von: Sheng, Zhang, et al.
Veröffentlicht: (2025)
Hardware-Agnostic and Insightful Efficiency Metrics for Accelerated Systems: Definition and Implementation within TALP
von: Rahimi, Ghazal, et al.
Veröffentlicht: (2026)
von: Rahimi, Ghazal, et al.
Veröffentlicht: (2026)
Mist: Efficient Distributed Training of Large Language Models via Memory-Parallelism Co-Optimization
von: Zhu, Zhanda, et al.
Veröffentlicht: (2025)
von: Zhu, Zhanda, et al.
Veröffentlicht: (2025)
AcceleratedKernels.jl: Cross-Architecture Parallel Algorithms from a Unified, Transpiled Codebase
von: Nicusan, Andrei-Leonard, et al.
Veröffentlicht: (2025)
von: Nicusan, Andrei-Leonard, et al.
Veröffentlicht: (2025)
Orthrus: Accelerating Multi-BFT Consensus through Concurrent Partial Ordering of Transactions (Extended Version)
von: Lyu, Hanzheng, et al.
Veröffentlicht: (2024)
von: Lyu, Hanzheng, et al.
Veröffentlicht: (2024)
Temporal Load Imbalance on Ondes3D Seismic Simulator for Different Multicore Architectures
von: Solórzano, Ana Luisa Veroneze, et al.
Veröffentlicht: (2024)
von: Solórzano, Ana Luisa Veroneze, et al.
Veröffentlicht: (2024)
Confidential Computing on NVIDIA Hopper GPUs: A Performance Benchmark Study
von: Zhu, Jianwei, et al.
Veröffentlicht: (2024)
von: Zhu, Jianwei, et al.
Veröffentlicht: (2024)
Tally: Non-Intrusive Performance Isolation for Concurrent Deep Learning Workloads
von: Zhao, Wei, et al.
Veröffentlicht: (2024)
von: Zhao, Wei, et al.
Veröffentlicht: (2024)
Black-Scholes Option Pricing on Intel CPUs and GPUs: Implementation on SYCL and Optimization Techniques
von: Panova, Elena, et al.
Veröffentlicht: (2022)
von: Panova, Elena, et al.
Veröffentlicht: (2022)
Ridgeline: A 2D Roofline Model for Distributed Systems
von: Checconi, Fabio, et al.
Veröffentlicht: (2022)
von: Checconi, Fabio, et al.
Veröffentlicht: (2022)
iSpLib: A Library for Accelerating Graph Neural Networks using Auto-tuned Sparse Operations
von: Anik, Md Saidul Hoque, et al.
Veröffentlicht: (2024)
von: Anik, Md Saidul Hoque, et al.
Veröffentlicht: (2024)
A Precision Emulation Approach to the GPU Acceleration of Ab Initio Electronic Structure Calculations
von: Liu, Hang, et al.
Veröffentlicht: (2026)
von: Liu, Hang, et al.
Veröffentlicht: (2026)
An Experimental Study of Different Aggregation Schemes in Semi-Asynchronous Federated Learning
von: Li, Yunbo, et al.
Veröffentlicht: (2024)
von: Li, Yunbo, et al.
Veröffentlicht: (2024)
PlantD: Performance, Latency ANalysis, and Testing for Data Pipelines -- An Open Source Measurement, Testing, and Simulation Framework
von: Bogart, Christopher, et al.
Veröffentlicht: (2025)
von: Bogart, Christopher, et al.
Veröffentlicht: (2025)
Fault-Tolerant Hybrid-Parallel Training at Scale with Reliable and Efficient In-memory Checkpointing
von: Wang, Yuxin, et al.
Veröffentlicht: (2023)
von: Wang, Yuxin, et al.
Veröffentlicht: (2023)
Accelerating Sparse Tensor Decomposition Using Adaptive Linearized Representation
von: Laukemann, Jan, et al.
Veröffentlicht: (2024)
von: Laukemann, Jan, et al.
Veröffentlicht: (2024)
On Orchestrating Parallel Broadcasts for Distributed Ledgers
von: Sheng, Peiyao, et al.
Veröffentlicht: (2024)
von: Sheng, Peiyao, et al.
Veröffentlicht: (2024)
Shifting the Sweet Spot: High-Performance Matrix-Free Method for High-Order Elasticity
von: Chang, Dali, et al.
Veröffentlicht: (2026)
von: Chang, Dali, et al.
Veröffentlicht: (2026)
HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices
von: Zhao, Xuanlei, et al.
Veröffentlicht: (2024)
von: Zhao, Xuanlei, et al.
Veröffentlicht: (2024)
Multi-DNN Inference of Sparse Models on Edge SoCs
von: Luo, Jiawei, et al.
Veröffentlicht: (2026)
von: Luo, Jiawei, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
DCC: Data-Centric Compilation of Machine Learning Kernels for Processing-In-Memory Architectures
von: Yang, Peiming, et al.
Veröffentlicht: (2025) -
SwiftFusion: Scalable Sequence Parallelism for Distributed Inference of Diffusion Transformers on GPUs
von: Yang, Jiacheng, et al.
Veröffentlicht: (2026) -
Spira: Exploiting Voxel Data Structural Properties for Efficient Sparse Convolution in Point Cloud Networks
von: Adamopoulos, Dionysios, et al.
Veröffentlicht: (2025) -
PyGim: An Efficient Graph Neural Network Library for Real Processing-In-Memory Architectures
von: Giannoula, Christina, et al.
Veröffentlicht: (2024) -
How to Rent GPUs on a Budget
von: Li, Zhouzi, et al.
Veröffentlicht: (2024)