Minimum Cost Loop Nests for Contraction of a Sparse Tensor with a Tensor Network
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kanakagiri, Raghavendra, Solomonik, Edgar |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Utilizing Sparsity in the GPU-accelerated Assembly of Schur Complement Matrices in Domain Decomposition Methods
von: Homola, Jakub, et al.
Veröffentlicht: (2025)
von: Homola, Jakub, et al.
Veröffentlicht: (2025)
Challenging Portability Paradigms: FPGA Acceleration Using SYCL and OpenCL
von: de Castro, Manuel, et al.
Veröffentlicht: (2024)
von: de Castro, Manuel, et al.
Veröffentlicht: (2024)
HYLU: Hybrid Parallel Sparse LU Factorization
von: Chen, Xiaoming
Veröffentlicht: (2025)
von: Chen, Xiaoming
Veröffentlicht: (2025)
Cascaded Prediction and Asynchronous Execution of Iterative Algorithms on Heterogeneous Platforms
von: Gao, Jianhua, et al.
Veröffentlicht: (2024)
von: Gao, Jianhua, et al.
Veröffentlicht: (2024)
Vectorized Adaptive Histograms for Sparse Oblique Forests
von: Lubonja, Ariel, et al.
Veröffentlicht: (2026)
von: Lubonja, Ariel, et al.
Veröffentlicht: (2026)
Stream parallel skeleton optimization
von: Aldinucci, Marco, et al.
Veröffentlicht: (2024)
von: Aldinucci, Marco, et al.
Veröffentlicht: (2024)
StreamFlow: cross-breeding cloud with HPC
von: Colonnelli, Iacopo, et al.
Veröffentlicht: (2020)
von: Colonnelli, Iacopo, et al.
Veröffentlicht: (2020)
Sharded Elimination and Combining for Highly-Efficient Concurrent Stacks
von: Singh, Ajay, et al.
Veröffentlicht: (2026)
von: Singh, Ajay, et al.
Veröffentlicht: (2026)
Optimizing Fine-Grained Parallelism Through Dynamic Load Balancing on Multi-Socket Many-Core Systems
von: Wang, Wenyi, et al.
Veröffentlicht: (2025)
von: Wang, Wenyi, et al.
Veröffentlicht: (2025)
AutoTSMM: An Auto-tuning Framework for Building High-Performance Tall-and-Skinny Matrix-Matrix Multiplication on CPUs
von: Li, Chendi, et al.
Veröffentlicht: (2022)
von: Li, Chendi, et al.
Veröffentlicht: (2022)
Enabling Practical Transparent Checkpointing for MPI: A Topological Sort Approach
von: Xu, Yao, et al.
Veröffentlicht: (2024)
von: Xu, Yao, et al.
Veröffentlicht: (2024)
NVLang: Unified Static Typing for Actor-Based Concurrency on the BEAM
von: Guerreiro, Miguel de Oliveira
Veröffentlicht: (2025)
von: Guerreiro, Miguel de Oliveira
Veröffentlicht: (2025)
VeriFx: Correct Replicated Data Types for the Masses
von: De Porre, Kevin, et al.
Veröffentlicht: (2022)
von: De Porre, Kevin, et al.
Veröffentlicht: (2022)
Canonicalization of Batched Einstein Summations for Tuning Retrieval
von: Kulkarni, Kaushik, et al.
Veröffentlicht: (2026)
von: Kulkarni, Kaushik, et al.
Veröffentlicht: (2026)
Solving Large Rank-Deficient Linear Least-Squares Problems on Shared-Memory CPU Architectures and GPU Architectures
von: Chillarón, Mónica, et al.
Veröffentlicht: (2024)
von: Chillarón, Mónica, et al.
Veröffentlicht: (2024)
A C++17 Thread Pool for High-Performance Scientific Computing
von: Shoshany, Barak
Veröffentlicht: (2021)
von: Shoshany, Barak
Veröffentlicht: (2021)
A Systematic Literature Survey of Sparse Matrix-Vector Multiplication
von: Gao, Jianhua, et al.
Veröffentlicht: (2024)
von: Gao, Jianhua, et al.
Veröffentlicht: (2024)
Fancy Some Chips for Your TeaStore? Modeling the Control of an Adaptable Discrete System
von: Gallone, Anna, et al.
Veröffentlicht: (2025)
von: Gallone, Anna, et al.
Veröffentlicht: (2025)
NM-SpMM: Accelerating Matrix Multiplication Using N:M Sparsity with GPGPU
von: Ma, Cong, et al.
Veröffentlicht: (2025)
von: Ma, Cong, et al.
Veröffentlicht: (2025)
High-Performance Tensor Contraction without Transposition
von: Matthews, Devin A.
Veröffentlicht: (2016)
von: Matthews, Devin A.
Veröffentlicht: (2016)
Rust vs. C for Python Libraries: Evaluating Rust-Compatible Bindings Toolchains
von: Amaral, Isabella Basso do, et al.
Veröffentlicht: (2025)
von: Amaral, Isabella Basso do, et al.
Veröffentlicht: (2025)
Hybrid Quantum-HPC Middleware Systems for Adaptive Resource, Workload and Task Management
von: Mantha, Pradeep, et al.
Veröffentlicht: (2026)
von: Mantha, Pradeep, et al.
Veröffentlicht: (2026)
SpaDA: A Spatial Dataflow Architecture Programming Language
von: Gianinazzi, Lukas, et al.
Veröffentlicht: (2025)
von: Gianinazzi, Lukas, et al.
Veröffentlicht: (2025)
DNA sequence alignment: An assignment for OpenMP, MPI, and CUDA/OpenCL
von: Gonzalez-Escribano, Arturo, et al.
Veröffentlicht: (2024)
von: Gonzalez-Escribano, Arturo, et al.
Veröffentlicht: (2024)
On Advanced Monte Carlo Methods for Linear Algebra on Advanced Accelerator Architectures
von: Lebedev, Anton, et al.
Veröffentlicht: (2024)
von: Lebedev, Anton, et al.
Veröffentlicht: (2024)
Categorical Message Passing Language (CaMPL) for programmers
von: Hashimoto, Daniel Kiyoshi, et al.
Veröffentlicht: (2026)
von: Hashimoto, Daniel Kiyoshi, et al.
Veröffentlicht: (2026)
Mapping Sparse Triangular Solves to GPUs via Fine-grained Domain Decomposition
von: Gondhalekar, Atharva, et al.
Veröffentlicht: (2025)
von: Gondhalekar, Atharva, et al.
Veröffentlicht: (2025)
Guess-Verify-Refine: Data-Aware Top-K for Sparse-Attention Decoding on Blackwell via Temporal Correlation
von: Cheng, Long, et al.
Veröffentlicht: (2026)
von: Cheng, Long, et al.
Veröffentlicht: (2026)
Assembly of FETI dual operator using CUDA
von: Homola, Jakub, et al.
Veröffentlicht: (2025)
von: Homola, Jakub, et al.
Veröffentlicht: (2025)
A Study of Performance Portability in Plasma Physics Simulations
von: Ruzicka, Josef, et al.
Veröffentlicht: (2024)
von: Ruzicka, Josef, et al.
Veröffentlicht: (2024)
cfdSCOPE: A Fluid-Dynamics Proxy App for Teaching Performance Engineering
von: Arzt, Peter, et al.
Veröffentlicht: (2025)
von: Arzt, Peter, et al.
Veröffentlicht: (2025)
Accelerating High-Order Finite Element Simulations at Extreme Scale with FP64 Tensor Cores
von: Tu, Jiqun, et al.
Veröffentlicht: (2026)
von: Tu, Jiqun, et al.
Veröffentlicht: (2026)
Dynamic Memory Management on GPUs with SYCL
von: Standish, Russell K.
Veröffentlicht: (2025)
von: Standish, Russell K.
Veröffentlicht: (2025)
Joint Training on AMD and NVIDIA GPUs
von: Hu, Jon, et al.
Veröffentlicht: (2026)
von: Hu, Jon, et al.
Veröffentlicht: (2026)
A Lock-Free, Fully GPU-Resident Architecture for the Verification of Goldbach's Conjecture
von: Llorente-Saguer, Isaac
Veröffentlicht: (2026)
von: Llorente-Saguer, Isaac
Veröffentlicht: (2026)
Communication-Efficient, 2D Parallel Stochastic Gradient Descent for Distributed-Memory Optimization
von: Devarakonda, Aditya, et al.
Veröffentlicht: (2025)
von: Devarakonda, Aditya, et al.
Veröffentlicht: (2025)
Distributed and heterogeneous tensor-vector contraction algorithms for high performance computing
von: Martinez-Ferrer, Pedro J., et al.
Veröffentlicht: (2025)
von: Martinez-Ferrer, Pedro J., et al.
Veröffentlicht: (2025)
How to Relax Instantly: Elastic Relaxation of Concurrent Data Structures
von: von Geijer, Kåre, et al.
Veröffentlicht: (2024)
von: von Geijer, Kåre, et al.
Veröffentlicht: (2024)
Speed, power and cost implications for GPU acceleration of Computational Fluid Dynamics on HPC systems
von: Cooper-Baldock, Zachary, et al.
Veröffentlicht: (2024)
von: Cooper-Baldock, Zachary, et al.
Veröffentlicht: (2024)
TC-GS: A Faster Gaussian Splatting Module Utilizing Tensor Cores
von: Liao, Zimu, et al.
Veröffentlicht: (2025)
von: Liao, Zimu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Utilizing Sparsity in the GPU-accelerated Assembly of Schur Complement Matrices in Domain Decomposition Methods
von: Homola, Jakub, et al.
Veröffentlicht: (2025) -
Challenging Portability Paradigms: FPGA Acceleration Using SYCL and OpenCL
von: de Castro, Manuel, et al.
Veröffentlicht: (2024) -
HYLU: Hybrid Parallel Sparse LU Factorization
von: Chen, Xiaoming
Veröffentlicht: (2025) -
Cascaded Prediction and Asynchronous Execution of Iterative Algorithms on Heterogeneous Platforms
von: Gao, Jianhua, et al.
Veröffentlicht: (2024) -
Vectorized Adaptive Histograms for Sparse Oblique Forests
von: Lubonja, Ariel, et al.
Veröffentlicht: (2026)