Do We Need Tensor Cores for Stencil Computations?
Fuente:
arXiv
Saved in:
| Main Authors: | Gu, Qiqi, Wu, Chenpeng, Shi, Heng, Yao, Jianguo, Guan, Haibing |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SPIDER: Unleashing Sparse Tensor Cores for Stencil Computation via Strided Swapping
by: GU, Qiqi, et al.
Published: (2025)
by: GU, Qiqi, et al.
Published: (2025)
Samoyeds: Accelerating MoE Models with Structured Sparsity Leveraging Sparse Tensor Cores
by: Wu, Chenpeng, et al.
Published: (2025)
by: Wu, Chenpeng, et al.
Published: (2025)
Stencil Matrixization
by: Zhao, Wenxuan, et al.
Published: (2023)
by: Zhao, Wenxuan, et al.
Published: (2023)
Evaluation of Programming Models and Performance for Stencil Computation on Current GPU Architectures
by: Shan, Baodi, et al.
Published: (2024)
by: Shan, Baodi, et al.
Published: (2024)
A Portable Framework for Accelerating Stencil Computations on Modern Node Architectures
by: Sai, Ryuichi, et al.
Published: (2023)
by: Sai, Ryuichi, et al.
Published: (2023)
Stencil Computations on Tenstorrent Wormhole
by: Piarulli, Lorenzo, et al.
Published: (2026)
by: Piarulli, Lorenzo, et al.
Published: (2026)
Persistent and Partitioned MPI for Stencil Communication
by: Collom, Gerald, et al.
Published: (2025)
by: Collom, Gerald, et al.
Published: (2025)
An Adaptive Distributed Stencil Abstraction for GPUs
by: Bhosale, Aditya, et al.
Published: (2025)
by: Bhosale, Aditya, et al.
Published: (2025)
To Repair or Not to Repair: Assessing Fault Resilience in MPI Stencil Applications
by: Rocco, Roberto, et al.
Published: (2024)
by: Rocco, Roberto, et al.
Published: (2024)
High Performance Unstructured SpMM Computation Using Tensor Cores
by: Okanovic, Patrik, et al.
Published: (2024)
by: Okanovic, Patrik, et al.
Published: (2024)
MMStencil: Optimizing High-order Stencils on Multicore CPU using Matrix Unit
by: Wang, Yinuo, et al.
Published: (2025)
by: Wang, Yinuo, et al.
Published: (2025)
QPU Micro-Kernels for Stencil Computation
by: Markidis, Stefano, et al.
Published: (2025)
by: Markidis, Stefano, et al.
Published: (2025)
Cyclic Data Streaming on GPUs for Short Range Stencils Applied to Molecular Dynamics
by: Rose, Martin, et al.
Published: (2025)
by: Rose, Martin, et al.
Published: (2025)
Accelerating Drug Discovery in AutoDock-GPU with Tensor Cores
by: Schieffer, Gabin, et al.
Published: (2024)
by: Schieffer, Gabin, et al.
Published: (2024)
Why Ethereum Needs Fairness Mechanisms that Do Not Depend on Participant Altruism
by: Spiesberger, Patrick, et al.
Published: (2026)
by: Spiesberger, Patrick, et al.
Published: (2026)
FedCod: An Efficient Communication Protocol for Cross-Silo Federated Learning with Coding
by: Yan, Peishen, et al.
Published: (2024)
by: Yan, Peishen, et al.
Published: (2024)
PUSHtap: PIM-based In-Memory HTAP with Unified Data Storage Format
by: Zhao, Yilong, et al.
Published: (2025)
by: Zhao, Yilong, et al.
Published: (2025)
Amoeba: Runtime Tensor Parallel Transformation for LLM Inference Services
by: Chen, Haoyu, et al.
Published: (2025)
by: Chen, Haoyu, et al.
Published: (2025)
Predictive Performance of Photonic SRAM-based In-Memory Computing for Tensor Decomposition
by: Wijeratne, Sasindu, et al.
Published: (2025)
by: Wijeratne, Sasindu, et al.
Published: (2025)
Minimizing Communication for Parallel Symmetric Tensor Times Same Vector Computation
by: Daas, Hussam Al, et al.
Published: (2025)
by: Daas, Hussam Al, et al.
Published: (2025)
Stencil Computations on Cerebras Wafer-Scale Engine
by: Belli, Elia, et al.
Published: (2026)
by: Belli, Elia, et al.
Published: (2026)
What Every Computer Scientist Needs To Know About Parallelization
by: Adefemi, Temitayo
Published: (2025)
by: Adefemi, Temitayo
Published: (2025)
Guaranteed DGEMM Accuracy While Using Reduced Precision Tensor Cores Through Extensions of the Ozaki Scheme
by: Schwarz, Angelika, et al.
Published: (2025)
by: Schwarz, Angelika, et al.
Published: (2025)
Can Tensor Cores Benefit Memory-Bound Kernels? (No!)
by: Zhang, Lingqi, et al.
Published: (2025)
by: Zhang, Lingqi, et al.
Published: (2025)
Accelerating Heterogeneous Tensor Parallelism via Flexible Workload Control
by: Wang, Zhigang, et al.
Published: (2024)
by: Wang, Zhigang, et al.
Published: (2024)
Navigating the Energy Doldrums: Can We Exploit Energy-Price Volatility To Lower the Cost of Computing?
by: Arzt, Peter, et al.
Published: (2025)
by: Arzt, Peter, et al.
Published: (2025)
cuFastTuckerPlus: A Stochastic Parallel Sparse FastTucker Decomposition Using GPU Tensor Cores
by: Li, Zixuan, et al.
Published: (2024)
by: Li, Zixuan, et al.
Published: (2024)
Do GPUs Really Need New Tabular File Formats?
by: Luo, Jigao, et al.
Published: (2026)
by: Luo, Jigao, et al.
Published: (2026)
Ray Tracing Cores for General-Purpose Computing: A Literature Review
by: Meneses, Enzo, et al.
Published: (2026)
by: Meneses, Enzo, et al.
Published: (2026)
FlashFuser: Expanding the Scale of Kernel Fusion for Compute-Intensive Operators via Inter-Core Connection
by: Huang, Ziyu, et al.
Published: (2025)
by: Huang, Ziyu, et al.
Published: (2025)
Low-ordered Orthogonal Voxel Finite Element with INT8 Tensor Cores for GPU-based Explicit Elastic Wave Propagation Analysis
by: Ichimura, Tsuyoshi, et al.
Published: (2024)
by: Ichimura, Tsuyoshi, et al.
Published: (2024)
NeutronTP: Load-Balanced Distributed Full-Graph GNN Training with Tensor Parallelism
by: Ai, Xin, et al.
Published: (2024)
by: Ai, Xin, et al.
Published: (2024)
HadaCore: Tensor Core Accelerated Hadamard Transform Kernel
by: Agarwal, Krish, et al.
Published: (2024)
by: Agarwal, Krish, et al.
Published: (2024)
An MLIR Lowering Pipeline for Stencils at Wafer-Scale
by: Stawinoga, Nicolai, et al.
Published: (2026)
by: Stawinoga, Nicolai, et al.
Published: (2026)
FWeb3: A Practical Incentive-Aware Federated Learning Framework
by: Yan, Peishen, et al.
Published: (2026)
by: Yan, Peishen, et al.
Published: (2026)
Synergistic Tensor and Pipeline Parallelism
by: Qi, Mengshi, et al.
Published: (2025)
by: Qi, Mengshi, et al.
Published: (2025)
A General Input-Dependent Colorless Computability Theorem and Applications to Core-Dependent Adversaries
by: Coutouly, Yannis, et al.
Published: (2025)
by: Coutouly, Yannis, et al.
Published: (2025)
Fast and Scalable Mixed Precision Euclidean Distance Calculations Using GPU Tensor Cores
by: Curless, Brian, et al.
Published: (2025)
by: Curless, Brian, et al.
Published: (2025)
Data-Centric Design: Introducing An Informatics Domain Model And Core Data Ontology For Computational Systems
by: Knowles, Paul, et al.
Published: (2024)
by: Knowles, Paul, et al.
Published: (2024)
CAT: Cellular Automata on Tensor cores
by: Navarro, Cristóbal A., et al.
Published: (2024)
by: Navarro, Cristóbal A., et al.
Published: (2024)
Similar Items
-
SPIDER: Unleashing Sparse Tensor Cores for Stencil Computation via Strided Swapping
by: GU, Qiqi, et al.
Published: (2025) -
Samoyeds: Accelerating MoE Models with Structured Sparsity Leveraging Sparse Tensor Cores
by: Wu, Chenpeng, et al.
Published: (2025) -
Stencil Matrixization
by: Zhao, Wenxuan, et al.
Published: (2023) -
Evaluation of Programming Models and Performance for Stencil Computation on Current GPU Architectures
by: Shan, Baodi, et al.
Published: (2024) -
A Portable Framework for Accelerating Stencil Computations on Modern Node Architectures
by: Sai, Ryuichi, et al.
Published: (2023)