Insum: Sparse GPU Kernels Simplified and Optimized with Indirect Einsums
Fuente:
arXiv
Saved in:
| Main Authors: | Won, Jaeyeon, Ahrens, Willow, Emer, Joel S., Amarasinghe, Saman |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Continuous Tensor Abstraction: Where Indices are Real
by: Won, Jaeyeon, et al.
Published: (2024)
by: Won, Jaeyeon, et al.
Published: (2024)
Mechanised Hypersafety Proofs about Structured Data: Extended Version
by: Gladshtein, Vladimir, et al.
Published: (2024)
by: Gladshtein, Vladimir, et al.
Published: (2024)
Flex Attention: A Programming Model for Generating Optimized Attention Kernels
by: Dong, Juechu, et al.
Published: (2024)
by: Dong, Juechu, et al.
Published: (2024)
Galley: Modern Query Optimization for Sparse Tensor Programs
by: Deeds, Kyle, et al.
Published: (2024)
by: Deeds, Kyle, et al.
Published: (2024)
SySTeC: A Symmetric Sparse Tensor Compiler
by: Patel, Radha, et al.
Published: (2024)
by: Patel, Radha, et al.
Published: (2024)
Optimizing Layout of Recursive Datatypes with Marmoset
by: Singhal, Vidush, et al.
Published: (2024)
by: Singhal, Vidush, et al.
Published: (2024)
Evaluating Compiler Optimization Impacts on zkVM Performance
by: Gassmann, Thomas, et al.
Published: (2025)
by: Gassmann, Thomas, et al.
Published: (2025)
Annotation-guided AoS-to-SoA conversions and GPU offloading with data views in C++
by: Radtke, Pawel K., et al.
Published: (2025)
by: Radtke, Pawel K., et al.
Published: (2025)
AutoLALA: Automatic Loop Algebraic Locality Analysis for AI and HPC Kernels
by: Zhu, Yifan, et al.
Published: (2026)
by: Zhu, Yifan, et al.
Published: (2026)
Stencil-Lifting: Hierarchical Recursive Lifting System for Extracting Summary of Stencil Kernel in Legacy Codes
by: Li, Mingyi, et al.
Published: (2025)
by: Li, Mingyi, et al.
Published: (2025)
The Next 700 ML-Enabled Compiler Optimizations
by: VenkataKeerthy, S., et al.
Published: (2023)
by: VenkataKeerthy, S., et al.
Published: (2023)
Rule-Based Graph Programs Matching the Time Complexity of Imperative Algorithms
by: Alaoui, Ziad Ismaili, et al.
Published: (2025)
by: Alaoui, Ziad Ismaili, et al.
Published: (2025)
It's Not Easy Being Green: On the Energy Efficiency of Programming Languages
by: van Kempen, Nicolas, et al.
Published: (2024)
by: van Kempen, Nicolas, et al.
Published: (2024)
CoNST: Code Generator for Sparse Tensor Networks
by: Raje, Saurabh, et al.
Published: (2024)
by: Raje, Saurabh, et al.
Published: (2024)
LOOPerSet: A Large-Scale Dataset for Data-Driven Polyhedral Compiler Optimization
by: Merouani, Massinissa, et al.
Published: (2025)
by: Merouani, Massinissa, et al.
Published: (2025)
Unified schemes for directive-based GPU offloading
by: Miki, Yohei, et al.
Published: (2024)
by: Miki, Yohei, et al.
Published: (2024)
Backwards Data-Flow Analysis using Prophecy Variables in the BuildIt System
by: Brahmakshatriya, Ajay, et al.
Published: (2026)
by: Brahmakshatriya, Ajay, et al.
Published: (2026)
ECCO: Evidence-Driven Causal Reasoning for Compiler Optimization
by: Pan, Haolin, et al.
Published: (2026)
by: Pan, Haolin, et al.
Published: (2026)
An Empirical Study on the Performance and Energy Usage of Compiled Python Code
by: Stoico, Vincenzo, et al.
Published: (2025)
by: Stoico, Vincenzo, et al.
Published: (2025)
Tidying Up the Address Space
by: Banakar, Vinay, et al.
Published: (2025)
by: Banakar, Vinay, et al.
Published: (2025)
Building an Accelerated OpenFOAM Proof-of-Concept Application using Modern C++
by: Malenza, Giulio, et al.
Published: (2025)
by: Malenza, Giulio, et al.
Published: (2025)
Latency Based Tiling
by: Cashman, Jack
Published: (2025)
by: Cashman, Jack
Published: (2025)
DaCe AD: Unifying High-Performance Automatic Differentiation for Machine Learning and Scientific Computing
by: Boudaoud, Afif, et al.
Published: (2025)
by: Boudaoud, Afif, et al.
Published: (2025)
Input-Gen: Guided Generation of Stateful Inputs for Testing, Tuning, and Training
by: Ivanov, Ivan R., et al.
Published: (2024)
by: Ivanov, Ivan R., et al.
Published: (2024)
Integration of a systolic array based hardware accelerator into a DNN operator auto-tuning framework
by: Peccia, F. N., et al.
Published: (2022)
by: Peccia, F. N., et al.
Published: (2022)
Runtime Verification on Abstract Finite State Models
by: Jevitha, KP, et al.
Published: (2024)
by: Jevitha, KP, et al.
Published: (2024)
Testing the Unknown: A Framework for OpenMP Testing via Random Program Generation
by: Laguna, Ignacio, et al.
Published: (2024)
by: Laguna, Ignacio, et al.
Published: (2024)
MapReplay: Trace-Driven Benchmark Generation for Java HashMap
by: Schiavio, Filippo, et al.
Published: (2026)
by: Schiavio, Filippo, et al.
Published: (2026)
Optimization of Armv9 architecture general large language model inference performance based on Llama.cpp
by: Chen, Longhao, et al.
Published: (2024)
by: Chen, Longhao, et al.
Published: (2024)
Runtime Repeated Recursion Unfolding in CHR: A Just-In-Time Online Program Optimization Strategy That Can Achieve Super-Linear Speedup
by: Fruehwirth, Thom
Published: (2023)
by: Fruehwirth, Thom
Published: (2023)
AutoKernel: Autonomous GPU Kernel Optimization via Iterative Agent-Driven Search
by: Jaber, Jaber, et al.
Published: (2026)
by: Jaber, Jaber, et al.
Published: (2026)
SoCal: A Language for Memory-Layout Factorization of Recursive Datatypes
by: Singhal, Vidush, et al.
Published: (2026)
by: Singhal, Vidush, et al.
Published: (2026)
Agentic Auto-Scheduling: An Experimental Study of LLM-Guided Loop Optimization
by: Merouani, Massinissa, et al.
Published: (2025)
by: Merouani, Massinissa, et al.
Published: (2025)
LOOPRAG: Enhancing Loop Transformation Optimization with Retrieval-Augmented Large Language Models
by: Zhi, Yijie, et al.
Published: (2025)
by: Zhi, Yijie, et al.
Published: (2025)
TAPA: A Scalable Task-Parallel Dataflow Programming Framework for Modern FPGAs with Co-Optimization of HLS and Physical Design
by: Guo, Licheng, et al.
Published: (2022)
by: Guo, Licheng, et al.
Published: (2022)
Comparing Parallel Functional Array Languages: Programming and Performance
by: van Balen, David, et al.
Published: (2025)
by: van Balen, David, et al.
Published: (2025)
Who Wins the Race? (R Vs Python) - An Exploratory Study on Energy Consumption of Machine Learning Algorithms
by: Chattaraj, Rajrupa, et al.
Published: (2025)
by: Chattaraj, Rajrupa, et al.
Published: (2025)
Library Liberation: Competitive Performance Matmul Through Compiler-composed Nanokernels
by: Thangamani, Arun, et al.
Published: (2025)
by: Thangamani, Arun, et al.
Published: (2025)
LEGO: A Layout Expression Language for Code Generation of Hierarchical Mapping
by: Tavakkoli, Amir Mohammad, et al.
Published: (2025)
by: Tavakkoli, Amir Mohammad, et al.
Published: (2025)
Iterating Pointers: Enabling Static Analysis for Loop-based Pointers
by: Lepori, Andrea, et al.
Published: (2025)
by: Lepori, Andrea, et al.
Published: (2025)
Similar Items
-
The Continuous Tensor Abstraction: Where Indices are Real
by: Won, Jaeyeon, et al.
Published: (2024) -
Mechanised Hypersafety Proofs about Structured Data: Extended Version
by: Gladshtein, Vladimir, et al.
Published: (2024) -
Flex Attention: A Programming Model for Generating Optimized Attention Kernels
by: Dong, Juechu, et al.
Published: (2024) -
Galley: Modern Query Optimization for Sparse Tensor Programs
by: Deeds, Kyle, et al.
Published: (2024) -
SySTeC: A Symmetric Sparse Tensor Compiler
by: Patel, Radha, et al.
Published: (2024)