Task-Based Tensor Computations on Modern GPUs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yadav, Rohan, Garland, Michael, Aiken, Alex, Bauer, Michael |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Optimal Software Pipelining and Warp Specialization for Tensor Core GPUs
von: Soi, Rupanshu, et al.
Veröffentlicht: (2025)
von: Soi, Rupanshu, et al.
Veröffentlicht: (2025)
On the Duality of Task and Actor Programming Models
von: Yadav, Rohan, et al.
Veröffentlicht: (2025)
von: Yadav, Rohan, et al.
Veröffentlicht: (2025)
Automatic Tracing in Task-Based Runtime Systems
von: Yadav, Rohan, et al.
Veröffentlicht: (2024)
von: Yadav, Rohan, et al.
Veröffentlicht: (2024)
Composing Distributed Computations Through Task and Kernel Fusion
von: Yadav, Rohan, et al.
Veröffentlicht: (2024)
von: Yadav, Rohan, et al.
Veröffentlicht: (2024)
Mapple: A Domain-Specific Language for Mapping Distributed Programs
von: Wei, Anjiang, et al.
Veröffentlicht: (2025)
von: Wei, Anjiang, et al.
Veröffentlicht: (2025)
Partitioning Unstructured Sparse Tensor Algebra for Load-Balanced Parallel Execution
von: Chougule, Atharva, et al.
Veröffentlicht: (2026)
von: Chougule, Atharva, et al.
Veröffentlicht: (2026)
Tawa: Automatic Warp Specialization for Modern GPUs with Asynchronous References
von: Chen, Hongzheng, et al.
Veröffentlicht: (2025)
von: Chen, Hongzheng, et al.
Veröffentlicht: (2025)
Galley: Modern Query Optimization for Sparse Tensor Programs
von: Deeds, Kyle, et al.
Veröffentlicht: (2024)
von: Deeds, Kyle, et al.
Veröffentlicht: (2024)
Rewrite System Showdown: Stochastic Search vs. EqSat
von: Hong, Qiantan, et al.
Veröffentlicht: (2026)
von: Hong, Qiantan, et al.
Veröffentlicht: (2026)
Compressing Structured Tensor Algebra
von: Ghorbani, Mahdi, et al.
Veröffentlicht: (2024)
von: Ghorbani, Mahdi, et al.
Veröffentlicht: (2024)
Tensor Evolution: A Framework for Fast Evaluation of Tensor Computations using Recurrences
von: Absar, Javed, et al.
Veröffentlicht: (2025)
von: Absar, Javed, et al.
Veröffentlicht: (2025)
Scaling Worst-Case Optimal Datalog to GPUs
von: Sun, Yihao, et al.
Veröffentlicht: (2026)
von: Sun, Yihao, et al.
Veröffentlicht: (2026)
Equivalence Checking of ML GPU Kernels
von: Dubey, Kshitij, et al.
Veröffentlicht: (2025)
von: Dubey, Kshitij, et al.
Veröffentlicht: (2025)
A Novel Compiler Transformation for Fast Sparse Matrix Multiplication in GPUs
von: Albakri, Hossein, et al.
Veröffentlicht: (2025)
von: Albakri, Hossein, et al.
Veröffentlicht: (2025)
SparseAuto: An Auto-Scheduler for Sparse Tensor Computations Using Recursive Loop Nest Restructuring
von: Dias, Adhitha, et al.
Veröffentlicht: (2023)
von: Dias, Adhitha, et al.
Veröffentlicht: (2023)
LTL learning on GPUs
von: Valizadeh, Mojtaba, et al.
Veröffentlicht: (2024)
von: Valizadeh, Mojtaba, et al.
Veröffentlicht: (2024)
Neptune: Advanced ML Operator Fusion for Locality and Parallelism on GPUs
von: Zhao, Yifan, et al.
Veröffentlicht: (2025)
von: Zhao, Yifan, et al.
Veröffentlicht: (2025)
Tenspiler: A Verified Lifting-Based Compiler for Tensor Operations (Extended Version)
von: Qiu, Jie, et al.
Veröffentlicht: (2024)
von: Qiu, Jie, et al.
Veröffentlicht: (2024)
Quokka: Accelerating Program Verification with LLMs via Invariant Synthesis
von: Wei, Anjiang, et al.
Veröffentlicht: (2025)
von: Wei, Anjiang, et al.
Veröffentlicht: (2025)
Domain-Specific Tensor Languages
von: Bernardy, Jean-Philippe, et al.
Veröffentlicht: (2023)
von: Bernardy, Jean-Philippe, et al.
Veröffentlicht: (2023)
Testing GPU Numerics: Finding Numerical Differences Between NVIDIA and AMD GPUs
von: Zahid, Anwar Hossain, et al.
Veröffentlicht: (2024)
von: Zahid, Anwar Hossain, et al.
Veröffentlicht: (2024)
Tensor Network Structure Search with Program Synthesis
von: Guo, Zheng, et al.
Veröffentlicht: (2025)
von: Guo, Zheng, et al.
Veröffentlicht: (2025)
The Continuous Tensor Abstraction: Where Indices are Real
von: Won, Jaeyeon, et al.
Veröffentlicht: (2024)
von: Won, Jaeyeon, et al.
Veröffentlicht: (2024)
Modernizing SMT-Based Type Error Localization
von: Kopinsky, Max, et al.
Veröffentlicht: (2024)
von: Kopinsky, Max, et al.
Veröffentlicht: (2024)
The T-Complexity Costs of Error Correction for Control Flow in Quantum Computation
von: Yuan, Charles, et al.
Veröffentlicht: (2023)
von: Yuan, Charles, et al.
Veröffentlicht: (2023)
DVM: A Bytecode Virtual Machine Approach for Dynamic Tensor Computation
von: Fang, Jingzhi, et al.
Veröffentlicht: (2026)
von: Fang, Jingzhi, et al.
Veröffentlicht: (2026)
SuperCoder: Assembly Program Superoptimization with Large Language Models
von: Wei, Anjiang, et al.
Veröffentlicht: (2025)
von: Wei, Anjiang, et al.
Veröffentlicht: (2025)
DITRON: Distributed Multi-level Tiling Compiler for Parallel Tensor Programs
von: Zheng, Size, et al.
Veröffentlicht: (2026)
von: Zheng, Size, et al.
Veröffentlicht: (2026)
A Few Fit Most: Improving Performance Portability of SGEMM on GPUs using Multi-Versioning
von: Hochgraf, Robert, et al.
Veröffentlicht: (2025)
von: Hochgraf, Robert, et al.
Veröffentlicht: (2025)
Pushing Tensor Accelerators Beyond MatMul in a User-Schedulable Language
von: Zhang, Yihong, et al.
Veröffentlicht: (2025)
von: Zhang, Yihong, et al.
Veröffentlicht: (2025)
Validation of Modern JSON Schema: Formalization and Complexity
von: Attouche, Lyes, et al.
Veröffentlicht: (2023)
von: Attouche, Lyes, et al.
Veröffentlicht: (2023)
Evaluating the Language-Based Security for Plugin Development
von: Liang, Naisheng, et al.
Veröffentlicht: (2024)
von: Liang, Naisheng, et al.
Veröffentlicht: (2024)
Elimination of annotation dependencies in validation for Modern JSON Schema
von: Attouche, Lyes, et al.
Veröffentlicht: (2025)
von: Attouche, Lyes, et al.
Veröffentlicht: (2025)
Fair intersection of seekable iterators
von: Arntzenius, Michael
Veröffentlicht: (2025)
von: Arntzenius, Michael
Veröffentlicht: (2025)
CodeARC: Benchmarking Reasoning Capabilities of LLM Agents for Inductive Program Synthesis
von: Wei, Anjiang, et al.
Veröffentlicht: (2025)
von: Wei, Anjiang, et al.
Veröffentlicht: (2025)
Learning Task Decomposition to Assist Humans in Competitive Programming
von: Wen, Jiaxin, et al.
Veröffentlicht: (2024)
von: Wen, Jiaxin, et al.
Veröffentlicht: (2024)
Strong Priority and Determinacy in Timed CCS
von: Liquori, Luigi, et al.
Veröffentlicht: (2024)
von: Liquori, Luigi, et al.
Veröffentlicht: (2024)
TENSURE: Fuzzing Sparse Tensor Compilers (Registered Report)
von: Mahathevan, Kabilan, et al.
Veröffentlicht: (2026)
von: Mahathevan, Kabilan, et al.
Veröffentlicht: (2026)
Portability of Fortran's `do concurrent' on GPUs
von: Caplan, Ronald M., et al.
Veröffentlicht: (2024)
von: Caplan, Ronald M., et al.
Veröffentlicht: (2024)
PSM: Policy Synchronised Deterministic Memory
von: Mendler, Michael, et al.
Veröffentlicht: (2025)
von: Mendler, Michael, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Optimal Software Pipelining and Warp Specialization for Tensor Core GPUs
von: Soi, Rupanshu, et al.
Veröffentlicht: (2025) -
On the Duality of Task and Actor Programming Models
von: Yadav, Rohan, et al.
Veröffentlicht: (2025) -
Automatic Tracing in Task-Based Runtime Systems
von: Yadav, Rohan, et al.
Veröffentlicht: (2024) -
Composing Distributed Computations Through Task and Kernel Fusion
von: Yadav, Rohan, et al.
Veröffentlicht: (2024) -
Mapple: A Domain-Specific Language for Mapping Distributed Programs
von: Wei, Anjiang, et al.
Veröffentlicht: (2025)