Accelerating Sparse MTTKRP for Small Tensor Decomposition on GPU
Fuente:
arXiv
Saved in:
| Main Authors: | Wijeratne, Sasindu, Kannan, Rajgopal, Prasanna, Viktor |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sparse MTTKRP Acceleration for Tensor Decomposition on GPU
by: Wijeratne, Sasindu, et al.
Published: (2024)
by: Wijeratne, Sasindu, et al.
Published: (2024)
AMPED: Accelerating MTTKRP for Billion-Scale Sparse Tensor Decomposition on Multiple GPUs
by: Wijeratne, Sasindu, et al.
Published: (2025)
by: Wijeratne, Sasindu, et al.
Published: (2025)
PRISM: Processing-In-Memory Sparse MTTKRP for Tensor Decomposition Acceleration
by: Pacheco, Daniel, et al.
Published: (2026)
by: Pacheco, Daniel, et al.
Published: (2026)
Predictive Performance of Photonic SRAM-based In-Memory Computing for Tensor Decomposition
by: Wijeratne, Sasindu, et al.
Published: (2025)
by: Wijeratne, Sasindu, et al.
Published: (2025)
System-Level Performance Modeling of Photonic In-Memory Computing
by: Arockiaraj, Jebacyril, et al.
Published: (2026)
by: Arockiaraj, Jebacyril, et al.
Published: (2026)
GCV-Turbo: End-to-end Acceleration of GNN-based Computer Vision Tasks on FPGA
by: Zhang, Bingyi, et al.
Published: (2024)
by: Zhang, Bingyi, et al.
Published: (2024)
A Unified CPU-GPU Protocol for GNN Training
by: Lin, Yi-Chien, et al.
Published: (2024)
by: Lin, Yi-Chien, et al.
Published: (2024)
Accelerating ViT Inference on FPGA through Static and Dynamic Pruning
by: Parikh, Dhruv, et al.
Published: (2024)
by: Parikh, Dhruv, et al.
Published: (2024)
ClusterViG: Efficient Globally Aware Vision GNNs via Image Partitioning
by: Parikh, Dhruv, et al.
Published: (2025)
by: Parikh, Dhruv, et al.
Published: (2025)
GPU Accelerated Sparse Cholesky Factorization
by: Karsavuran, M. Ozan, et al.
Published: (2024)
by: Karsavuran, M. Ozan, et al.
Published: (2024)
cuFastTuckerPlus: A Stochastic Parallel Sparse FastTucker Decomposition Using GPU Tensor Cores
by: Li, Zixuan, et al.
Published: (2024)
by: Li, Zixuan, et al.
Published: (2024)
Accelerating Dynamic Image Graph Construction on FPGA for Vision GNNs
by: Ramachandran, Anvitha, et al.
Published: (2025)
by: Ramachandran, Anvitha, et al.
Published: (2025)
Accelerating Drug Discovery in AutoDock-GPU with Tensor Cores
by: Schieffer, Gabin, et al.
Published: (2024)
by: Schieffer, Gabin, et al.
Published: (2024)
AsyncSparse: Accelerating Sparse Matrix-Matrix Multiplication on Asynchronous GPU Architectures
by: Liu, Jie, et al.
Published: (2026)
by: Liu, Jie, et al.
Published: (2026)
VTR: An Optimized Vision Transformer for SAR ATR Acceleration on FPGA
by: Wickramasinghe, Sachini, et al.
Published: (2024)
by: Wickramasinghe, Sachini, et al.
Published: (2024)
HC-SpMM: Accelerating Sparse Matrix-Matrix Multiplication for Graphs with Hybrid GPU Cores
by: Li, Zhonggen, et al.
Published: (2024)
by: Li, Zhonggen, et al.
Published: (2024)
Accelerating Biclique Counting on GPU
by: Qiu, Linshan, et al.
Published: (2024)
by: Qiu, Linshan, et al.
Published: (2024)
Accelerating Sparse Tensor Decomposition Using Adaptive Linearized Representation
by: Laukemann, Jan, et al.
Published: (2024)
by: Laukemann, Jan, et al.
Published: (2024)
A Scalable Recipe on SuperMUC-NG Phase 2: Efficient Large-Scale Training of Language Models
by: Rajgopal, Ajay Navilarekal, et al.
Published: (2026)
by: Rajgopal, Ajay Navilarekal, et al.
Published: (2026)
CRIUgpu: Transparent Checkpointing of GPU-Accelerated Workloads
by: Stoyanov, Radostin, et al.
Published: (2025)
by: Stoyanov, Radostin, et al.
Published: (2025)
GPU-Accelerated Batch-Dynamic Subgraph Matching
by: Qiu, Linshan, et al.
Published: (2024)
by: Qiu, Linshan, et al.
Published: (2024)
Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor Cores
by: Zhao, Haisha, et al.
Published: (2025)
by: Zhao, Haisha, et al.
Published: (2025)
Benchmarking the Performance of Large Language Models on the Cerebras Wafer Scale Engine
by: Zhang, Zuoning, et al.
Published: (2024)
by: Zhang, Zuoning, et al.
Published: (2024)
Federated Learning Using Coupled Tensor Train Decomposition
by: Zhang, Xiangtao, et al.
Published: (2024)
by: Zhang, Xiangtao, et al.
Published: (2024)
Accelerating Intra-Node GPU-to-GPU Communication Through Multi-Path Transfers with CUDA Graphs
by: Sojoodi, Amirhossein, et al.
Published: (2026)
by: Sojoodi, Amirhossein, et al.
Published: (2026)
GPZ: GPU-Accelerated Lossy Compressor for Particle Data
by: Li, Ruoyu, et al.
Published: (2025)
by: Li, Ruoyu, et al.
Published: (2025)
PICO: Accelerating All k-Core Paradigms on GPU
by: Zhao, Chen, et al.
Published: (2024)
by: Zhao, Chen, et al.
Published: (2024)
Efficient Accelerated Graph Edit Distance Computation on GPU
by: Dabah, Adel, et al.
Published: (2026)
by: Dabah, Adel, et al.
Published: (2026)
Speeding up Local Optimization in Vehicle Routing with Tensor-based GPU Acceleration
by: Lei, Zhenyu, et al.
Published: (2025)
by: Lei, Zhenyu, et al.
Published: (2025)
GPU-Accelerated Distributed QAOA on Large-scale HPC Ecosystems
by: Xu, Zhihao, et al.
Published: (2025)
by: Xu, Zhihao, et al.
Published: (2025)
PilotANN: Memory-Bounded GPU Acceleration for Vector Search
by: Gui, Yuntao, et al.
Published: (2025)
by: Gui, Yuntao, et al.
Published: (2025)
GPU Acceleration for Faster Evolutionary Spatial Cyclic Game Systems
by: Sinadjan, Louie
Published: (2025)
by: Sinadjan, Louie
Published: (2025)
A Preliminary Study on Accelerating Simulation Optimization with GPU Implementation
by: He, Jinghai, et al.
Published: (2024)
by: He, Jinghai, et al.
Published: (2024)
A Practical GPU-Accelerated Implementation of Orthogonal Matching Pursuit
by: Lubonja, Ariel, et al.
Published: (2024)
by: Lubonja, Ariel, et al.
Published: (2024)
SOLANET: Distributed Neighbor Graph Construction on GPU-Accelerated Systems
by: Iwabuchi, Keita, et al.
Published: (2026)
by: Iwabuchi, Keita, et al.
Published: (2026)
ScalableHD: Scalable and High-Throughput Hyperdimensional Computing Inference on Multi-Core CPUs
by: Parikh, Dhruv, et al.
Published: (2025)
by: Parikh, Dhruv, et al.
Published: (2025)
Collaborative Inference Acceleration with Non-Penetrative Tensor Partitioning
by: Liu, Zhibang, et al.
Published: (2025)
by: Liu, Zhibang, et al.
Published: (2025)
GraphLeap: Decoupling Graph Construction and Convolution for Vision GNN Acceleration on FPGA
by: Ramachandran, Anvitha, et al.
Published: (2026)
by: Ramachandran, Anvitha, et al.
Published: (2026)
GPU-Accelerated Modified Bessel Function of the Second Kind for Gaussian Processes
by: Geng, Zipei, et al.
Published: (2025)
by: Geng, Zipei, et al.
Published: (2025)
Dataflow-Oriented Classification and Performance Analysis of GPU-Accelerated Homomorphic Encryption
by: Nozaki, Ai, et al.
Published: (2026)
by: Nozaki, Ai, et al.
Published: (2026)
Similar Items
-
Sparse MTTKRP Acceleration for Tensor Decomposition on GPU
by: Wijeratne, Sasindu, et al.
Published: (2024) -
AMPED: Accelerating MTTKRP for Billion-Scale Sparse Tensor Decomposition on Multiple GPUs
by: Wijeratne, Sasindu, et al.
Published: (2025) -
PRISM: Processing-In-Memory Sparse MTTKRP for Tensor Decomposition Acceleration
by: Pacheco, Daniel, et al.
Published: (2026) -
Predictive Performance of Photonic SRAM-based In-Memory Computing for Tensor Decomposition
by: Wijeratne, Sasindu, et al.
Published: (2025) -
System-Level Performance Modeling of Photonic In-Memory Computing
by: Arockiaraj, Jebacyril, et al.
Published: (2026)