Task-Based Tensor Computations on Modern GPUs
Fuente:
arXiv
Saved in:
| Main Authors: | Yadav, Rohan, Garland, Michael, Aiken, Alex, Bauer, Michael |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Optimal Software Pipelining and Warp Specialization for Tensor Core GPUs
by: Soi, Rupanshu, et al.
Published: (2025)
by: Soi, Rupanshu, et al.
Published: (2025)
On the Duality of Task and Actor Programming Models
by: Yadav, Rohan, et al.
Published: (2025)
by: Yadav, Rohan, et al.
Published: (2025)
Automatic Tracing in Task-Based Runtime Systems
by: Yadav, Rohan, et al.
Published: (2024)
by: Yadav, Rohan, et al.
Published: (2024)
Composing Distributed Computations Through Task and Kernel Fusion
by: Yadav, Rohan, et al.
Published: (2024)
by: Yadav, Rohan, et al.
Published: (2024)
Mapple: A Domain-Specific Language for Mapping Distributed Programs
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
Partitioning Unstructured Sparse Tensor Algebra for Load-Balanced Parallel Execution
by: Chougule, Atharva, et al.
Published: (2026)
by: Chougule, Atharva, et al.
Published: (2026)
Tawa: Automatic Warp Specialization for Modern GPUs with Asynchronous References
by: Chen, Hongzheng, et al.
Published: (2025)
by: Chen, Hongzheng, et al.
Published: (2025)
Galley: Modern Query Optimization for Sparse Tensor Programs
by: Deeds, Kyle, et al.
Published: (2024)
by: Deeds, Kyle, et al.
Published: (2024)
Rewrite System Showdown: Stochastic Search vs. EqSat
by: Hong, Qiantan, et al.
Published: (2026)
by: Hong, Qiantan, et al.
Published: (2026)
Compressing Structured Tensor Algebra
by: Ghorbani, Mahdi, et al.
Published: (2024)
by: Ghorbani, Mahdi, et al.
Published: (2024)
Tensor Evolution: A Framework for Fast Evaluation of Tensor Computations using Recurrences
by: Absar, Javed, et al.
Published: (2025)
by: Absar, Javed, et al.
Published: (2025)
Scaling Worst-Case Optimal Datalog to GPUs
by: Sun, Yihao, et al.
Published: (2026)
by: Sun, Yihao, et al.
Published: (2026)
Equivalence Checking of ML GPU Kernels
by: Dubey, Kshitij, et al.
Published: (2025)
by: Dubey, Kshitij, et al.
Published: (2025)
A Novel Compiler Transformation for Fast Sparse Matrix Multiplication in GPUs
by: Albakri, Hossein, et al.
Published: (2025)
by: Albakri, Hossein, et al.
Published: (2025)
SparseAuto: An Auto-Scheduler for Sparse Tensor Computations Using Recursive Loop Nest Restructuring
by: Dias, Adhitha, et al.
Published: (2023)
by: Dias, Adhitha, et al.
Published: (2023)
LTL learning on GPUs
by: Valizadeh, Mojtaba, et al.
Published: (2024)
by: Valizadeh, Mojtaba, et al.
Published: (2024)
Neptune: Advanced ML Operator Fusion for Locality and Parallelism on GPUs
by: Zhao, Yifan, et al.
Published: (2025)
by: Zhao, Yifan, et al.
Published: (2025)
Tenspiler: A Verified Lifting-Based Compiler for Tensor Operations (Extended Version)
by: Qiu, Jie, et al.
Published: (2024)
by: Qiu, Jie, et al.
Published: (2024)
Quokka: Accelerating Program Verification with LLMs via Invariant Synthesis
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
Domain-Specific Tensor Languages
by: Bernardy, Jean-Philippe, et al.
Published: (2023)
by: Bernardy, Jean-Philippe, et al.
Published: (2023)
Testing GPU Numerics: Finding Numerical Differences Between NVIDIA and AMD GPUs
by: Zahid, Anwar Hossain, et al.
Published: (2024)
by: Zahid, Anwar Hossain, et al.
Published: (2024)
Tensor Network Structure Search with Program Synthesis
by: Guo, Zheng, et al.
Published: (2025)
by: Guo, Zheng, et al.
Published: (2025)
The Continuous Tensor Abstraction: Where Indices are Real
by: Won, Jaeyeon, et al.
Published: (2024)
by: Won, Jaeyeon, et al.
Published: (2024)
Modernizing SMT-Based Type Error Localization
by: Kopinsky, Max, et al.
Published: (2024)
by: Kopinsky, Max, et al.
Published: (2024)
The T-Complexity Costs of Error Correction for Control Flow in Quantum Computation
by: Yuan, Charles, et al.
Published: (2023)
by: Yuan, Charles, et al.
Published: (2023)
DVM: A Bytecode Virtual Machine Approach for Dynamic Tensor Computation
by: Fang, Jingzhi, et al.
Published: (2026)
by: Fang, Jingzhi, et al.
Published: (2026)
SuperCoder: Assembly Program Superoptimization with Large Language Models
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
DITRON: Distributed Multi-level Tiling Compiler for Parallel Tensor Programs
by: Zheng, Size, et al.
Published: (2026)
by: Zheng, Size, et al.
Published: (2026)
A Few Fit Most: Improving Performance Portability of SGEMM on GPUs using Multi-Versioning
by: Hochgraf, Robert, et al.
Published: (2025)
by: Hochgraf, Robert, et al.
Published: (2025)
Pushing Tensor Accelerators Beyond MatMul in a User-Schedulable Language
by: Zhang, Yihong, et al.
Published: (2025)
by: Zhang, Yihong, et al.
Published: (2025)
Validation of Modern JSON Schema: Formalization and Complexity
by: Attouche, Lyes, et al.
Published: (2023)
by: Attouche, Lyes, et al.
Published: (2023)
Evaluating the Language-Based Security for Plugin Development
by: Liang, Naisheng, et al.
Published: (2024)
by: Liang, Naisheng, et al.
Published: (2024)
Elimination of annotation dependencies in validation for Modern JSON Schema
by: Attouche, Lyes, et al.
Published: (2025)
by: Attouche, Lyes, et al.
Published: (2025)
Fair intersection of seekable iterators
by: Arntzenius, Michael
Published: (2025)
by: Arntzenius, Michael
Published: (2025)
CodeARC: Benchmarking Reasoning Capabilities of LLM Agents for Inductive Program Synthesis
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
Learning Task Decomposition to Assist Humans in Competitive Programming
by: Wen, Jiaxin, et al.
Published: (2024)
by: Wen, Jiaxin, et al.
Published: (2024)
Strong Priority and Determinacy in Timed CCS
by: Liquori, Luigi, et al.
Published: (2024)
by: Liquori, Luigi, et al.
Published: (2024)
TENSURE: Fuzzing Sparse Tensor Compilers (Registered Report)
by: Mahathevan, Kabilan, et al.
Published: (2026)
by: Mahathevan, Kabilan, et al.
Published: (2026)
Portability of Fortran's `do concurrent' on GPUs
by: Caplan, Ronald M., et al.
Published: (2024)
by: Caplan, Ronald M., et al.
Published: (2024)
PSM: Policy Synchronised Deterministic Memory
by: Mendler, Michael, et al.
Published: (2025)
by: Mendler, Michael, et al.
Published: (2025)
Similar Items
-
Optimal Software Pipelining and Warp Specialization for Tensor Core GPUs
by: Soi, Rupanshu, et al.
Published: (2025) -
On the Duality of Task and Actor Programming Models
by: Yadav, Rohan, et al.
Published: (2025) -
Automatic Tracing in Task-Based Runtime Systems
by: Yadav, Rohan, et al.
Published: (2024) -
Composing Distributed Computations Through Task and Kernel Fusion
by: Yadav, Rohan, et al.
Published: (2024) -
Mapple: A Domain-Specific Language for Mapping Distributed Programs
by: Wei, Anjiang, et al.
Published: (2025)