Low-ordered Orthogonal Voxel Finite Element with INT8 Tensor Cores for GPU-based Explicit Elastic Wave Propagation Analysis
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ichimura, Tsuyoshi, Fujita, Kohei, Hori, Muneo, Lalith, Maddegedara |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Accelerating Nonlinear Time-History Analysis with Complex Constitutive Laws via Heterogeneous Memory Management: From 3D Seismic Simulation to Neural Network Training
von: Ichimura, Tsuyoshi, et al.
Veröffentlicht: (2026)
von: Ichimura, Tsuyoshi, et al.
Veröffentlicht: (2026)
Accelerating Drug Discovery in AutoDock-GPU with Tensor Cores
von: Schieffer, Gabin, et al.
Veröffentlicht: (2024)
von: Schieffer, Gabin, et al.
Veröffentlicht: (2024)
A Practical GPU-Accelerated Implementation of Orthogonal Matching Pursuit
von: Lubonja, Ariel, et al.
Veröffentlicht: (2024)
von: Lubonja, Ariel, et al.
Veröffentlicht: (2024)
cuFastTuckerPlus: A Stochastic Parallel Sparse FastTucker Decomposition Using GPU Tensor Cores
von: Li, Zixuan, et al.
Veröffentlicht: (2024)
von: Li, Zixuan, et al.
Veröffentlicht: (2024)
Accelerating Sparse MTTKRP for Small Tensor Decomposition on GPU
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2025)
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2025)
Fast and Scalable Mixed Precision Euclidean Distance Calculations Using GPU Tensor Cores
von: Curless, Brian, et al.
Veröffentlicht: (2025)
von: Curless, Brian, et al.
Veröffentlicht: (2025)
Accelerating High-Order Finite Element Simulations at Extreme Scale with FP64 Tensor Cores
von: Tu, Jiqun, et al.
Veröffentlicht: (2026)
von: Tu, Jiqun, et al.
Veröffentlicht: (2026)
Dilu: Enabling GPU Resourcing-on-Demand for Serverless DL Serving via Introspective Elasticity
von: Lv, Cunchi, et al.
Veröffentlicht: (2025)
von: Lv, Cunchi, et al.
Veröffentlicht: (2025)
AnchorTP: Resilient LLM Inference with State-Preserving Elastic Tensor Parallelism
von: Xu, Wendong, et al.
Veröffentlicht: (2025)
von: Xu, Wendong, et al.
Veröffentlicht: (2025)
Convex Hull 3D Filtering with GPU Ray Tracing and Tensor Cores
von: Carrasco, Roberto, et al.
Veröffentlicht: (2026)
von: Carrasco, Roberto, et al.
Veröffentlicht: (2026)
High-Performance and Power-Efficient Emulation of Matrix Multiplication using INT8 Matrix Engines
von: Uchino, Yuki, et al.
Veröffentlicht: (2025)
von: Uchino, Yuki, et al.
Veröffentlicht: (2025)
PICO: Accelerating All k-Core Paradigms on GPU
von: Zhao, Chen, et al.
Veröffentlicht: (2024)
von: Zhao, Chen, et al.
Veröffentlicht: (2024)
Large Scale Multi-GPU Based Parallel Traffic Simulation for Accelerated Traffic Assignment and Propagation
von: Jiang, Xuan, et al.
Veröffentlicht: (2024)
von: Jiang, Xuan, et al.
Veröffentlicht: (2024)
ElasWave: An Elastic-Native System for Scalable Hybrid-Parallel Training
von: Kang, Xueze, et al.
Veröffentlicht: (2025)
von: Kang, Xueze, et al.
Veröffentlicht: (2025)
Do We Need Tensor Cores for Stencil Computations?
von: Gu, Qiqi, et al.
Veröffentlicht: (2026)
von: Gu, Qiqi, et al.
Veröffentlicht: (2026)
High Performance Unstructured SpMM Computation Using Tensor Cores
von: Okanovic, Patrik, et al.
Veröffentlicht: (2024)
von: Okanovic, Patrik, et al.
Veröffentlicht: (2024)
Sparse MTTKRP Acceleration for Tensor Decomposition on GPU
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2024)
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2024)
HC-SpMM: Accelerating Sparse Matrix-Matrix Multiplication for Graphs with Hybrid GPU Cores
von: Li, Zhonggen, et al.
Veröffentlicht: (2024)
von: Li, Zhonggen, et al.
Veröffentlicht: (2024)
Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient Inference
von: Yu, Minchen, et al.
Veröffentlicht: (2023)
von: Yu, Minchen, et al.
Veröffentlicht: (2023)
SPIDER: Unleashing Sparse Tensor Cores for Stencil Computation via Strided Swapping
von: GU, Qiqi, et al.
Veröffentlicht: (2025)
von: GU, Qiqi, et al.
Veröffentlicht: (2025)
Can Tensor Cores Benefit Memory-Bound Kernels? (No!)
von: Zhang, Lingqi, et al.
Veröffentlicht: (2025)
von: Zhang, Lingqi, et al.
Veröffentlicht: (2025)
Guaranteed DGEMM Accuracy While Using Reduced Precision Tensor Cores Through Extensions of the Ozaki Scheme
von: Schwarz, Angelika, et al.
Veröffentlicht: (2025)
von: Schwarz, Angelika, et al.
Veröffentlicht: (2025)
Improving GPU Multi-Tenancy Through Dynamic Multi-Instance GPU Reconfiguration
von: Wang, Tianyu, et al.
Veröffentlicht: (2024)
von: Wang, Tianyu, et al.
Veröffentlicht: (2024)
Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor Cores
von: Zhao, Haisha, et al.
Veröffentlicht: (2025)
von: Zhao, Haisha, et al.
Veröffentlicht: (2025)
TensorHub: Scalable and Elastic Weight Transfer for LLM RL Training
von: Ye, Chenhao, et al.
Veröffentlicht: (2026)
von: Ye, Chenhao, et al.
Veröffentlicht: (2026)
Accelerating Biclique Counting on GPU
von: Qiu, Linshan, et al.
Veröffentlicht: (2024)
von: Qiu, Linshan, et al.
Veröffentlicht: (2024)
GPU Sharing with Triples Mode
von: Byun, Chansup, et al.
Veröffentlicht: (2024)
von: Byun, Chansup, et al.
Veröffentlicht: (2024)
LoHan: Low-Cost High-Performance Framework to Fine-Tune 100B Model on a Consumer GPU
von: Liao, Changyue, et al.
Veröffentlicht: (2024)
von: Liao, Changyue, et al.
Veröffentlicht: (2024)
ParvaGPU: Efficient Spatial GPU Sharing for Large-Scale DNN Inference in Cloud Environments
von: Lee, Munkyu, et al.
Veröffentlicht: (2024)
von: Lee, Munkyu, et al.
Veröffentlicht: (2024)
Accelerating Intra-Node GPU-to-GPU Communication Through Multi-Path Transfers with CUDA Graphs
von: Sojoodi, Amirhossein, et al.
Veröffentlicht: (2026)
von: Sojoodi, Amirhossein, et al.
Veröffentlicht: (2026)
Speeding up Local Optimization in Vehicle Routing with Tensor-based GPU Acceleration
von: Lei, Zhenyu, et al.
Veröffentlicht: (2025)
von: Lei, Zhenyu, et al.
Veröffentlicht: (2025)
HadaCore: Tensor Core Accelerated Hadamard Transform Kernel
von: Agarwal, Krish, et al.
Veröffentlicht: (2024)
von: Agarwal, Krish, et al.
Veröffentlicht: (2024)
DuaLip-GPU Technical Report
von: Dexter, Gregory, et al.
Veröffentlicht: (2026)
von: Dexter, Gregory, et al.
Veröffentlicht: (2026)
Incidence Constraints in Hypergraph Partitioning on GPU
von: Ronzani, Marco, et al.
Veröffentlicht: (2026)
von: Ronzani, Marco, et al.
Veröffentlicht: (2026)
Predictable LLM Serving on GPU Clusters
von: Darzi, Erfan, et al.
Veröffentlicht: (2025)
von: Darzi, Erfan, et al.
Veröffentlicht: (2025)
GPU Accelerated Sparse Cholesky Factorization
von: Karsavuran, M. Ozan, et al.
Veröffentlicht: (2024)
von: Karsavuran, M. Ozan, et al.
Veröffentlicht: (2024)
Heat: Satellite's meat is GPU's poison
von: Yuan, Zhehu, et al.
Veröffentlicht: (2024)
von: Yuan, Zhehu, et al.
Veröffentlicht: (2024)
HAS-GPU: Efficient Hybrid Auto-scaling with Fine-grained GPU Allocation for SLO-aware Serverless Inferences
von: Gu, Jianfeng, et al.
Veröffentlicht: (2025)
von: Gu, Jianfeng, et al.
Veröffentlicht: (2025)
An Elastic Job Scheduler for HPC Applications on the Cloud
von: Bhosale, Aditya, et al.
Veröffentlicht: (2025)
von: Bhosale, Aditya, et al.
Veröffentlicht: (2025)
Enabling Elastic Model Serving with MultiWorld
von: Lee, Myungjin, et al.
Veröffentlicht: (2024)
von: Lee, Myungjin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Accelerating Nonlinear Time-History Analysis with Complex Constitutive Laws via Heterogeneous Memory Management: From 3D Seismic Simulation to Neural Network Training
von: Ichimura, Tsuyoshi, et al.
Veröffentlicht: (2026) -
Accelerating Drug Discovery in AutoDock-GPU with Tensor Cores
von: Schieffer, Gabin, et al.
Veröffentlicht: (2024) -
A Practical GPU-Accelerated Implementation of Orthogonal Matching Pursuit
von: Lubonja, Ariel, et al.
Veröffentlicht: (2024) -
cuFastTuckerPlus: A Stochastic Parallel Sparse FastTucker Decomposition Using GPU Tensor Cores
von: Li, Zixuan, et al.
Veröffentlicht: (2024) -
Accelerating Sparse MTTKRP for Small Tensor Decomposition on GPU
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2025)