tritonBLAS: Triton-based Analytical Approach for GEMM Kernel Parameter Selection
Fuente:
arXiv
Salvato in:
| Autori principali: | Swann, Ryan, Osama, Muhammad, Guo, Xiaohu, Nelson, Bryant, Zhang, Lixun, Brown, Alex, Ong, Yen, Yazdani, Ali, Siddens, Sean, Dasika, Ganesh, Underwood, Alex |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Fleet: Hierarchical Task-based Abstraction for Megakernels on Multi-Die GPUs
di: Chowdhary, Sangeeta, et al.
Pubblicazione: (2026)
di: Chowdhary, Sangeeta, et al.
Pubblicazione: (2026)
HipKittens: Fast and Furious AMD Kernels
di: Hu, William, et al.
Pubblicazione: (2025)
di: Hu, William, et al.
Pubblicazione: (2025)
SwizzlePerf: Hardware-Aware LLMs for GPU Kernel Performance Optimization
di: Tschand, Arya, et al.
Pubblicazione: (2025)
di: Tschand, Arya, et al.
Pubblicazione: (2025)
Seer: Predictive Runtime Kernel Selection for Irregular Problems
di: Swann, Ryan, et al.
Pubblicazione: (2024)
di: Swann, Ryan, et al.
Pubblicazione: (2024)
TritonForge: Profiling-Guided Framework for Automated Triton Kernel Optimization
di: Li, Haonan, et al.
Pubblicazione: (2025)
di: Li, Haonan, et al.
Pubblicazione: (2025)
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler
di: Zheng, Size, et al.
Pubblicazione: (2025)
di: Zheng, Size, et al.
Pubblicazione: (2025)
Liger Kernel: Efficient Triton Kernels for LLM Training
di: Hsu, Pin-Lun, et al.
Pubblicazione: (2024)
di: Hsu, Pin-Lun, et al.
Pubblicazione: (2024)
The Anatomy of a Triton Attention Kernel
di: Ringlein, Burkhard, et al.
Pubblicazione: (2025)
di: Ringlein, Burkhard, et al.
Pubblicazione: (2025)
Optimizing Attention on GPUs by Exploiting GPU Architectural NUMA Effects
di: Choudhary, Mansi, et al.
Pubblicazione: (2025)
di: Choudhary, Mansi, et al.
Pubblicazione: (2025)
Stream-K++: Adaptive GPU GEMM Kernel Scheduling and Selection using Bloom Filters
di: Sadasivan, Harisankar, et al.
Pubblicazione: (2024)
di: Sadasivan, Harisankar, et al.
Pubblicazione: (2024)
Iris: First-Class Multi-GPU Programming Experience in Triton
di: Awad, Muhammad, et al.
Pubblicazione: (2025)
di: Awad, Muhammad, et al.
Pubblicazione: (2025)
Dr. Kernel: Reinforcement Learning Done Right for Triton Kernel Generations
di: Liu, Wei, et al.
Pubblicazione: (2026)
di: Liu, Wei, et al.
Pubblicazione: (2026)
LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving
di: Hu, Huanqi, et al.
Pubblicazione: (2025)
di: Hu, Huanqi, et al.
Pubblicazione: (2025)
GUP Effective Metric Without GUP: Implications for the Sign of GUP Parameter and Quantum Bounce
di: Ong, Yen Chin
Pubblicazione: (2025)
di: Ong, Yen Chin
Pubblicazione: (2025)
Geak: Introducing Triton Kernel AI Agent & Evaluation Benchmarks
di: Wang, Jianghui, et al.
Pubblicazione: (2025)
di: Wang, Jianghui, et al.
Pubblicazione: (2025)
Eliminating Multi-GPU Performance Taxes: A Systems Approach to Efficient Distributed LLMs
di: Trifan, Octavian Alexandru, et al.
Pubblicazione: (2025)
di: Trifan, Octavian Alexandru, et al.
Pubblicazione: (2025)
Sparton: Fast and Memory-Efficient Triton Kernel for Learned Sparse Retrieval
di: Nguyen, Thong, et al.
Pubblicazione: (2026)
di: Nguyen, Thong, et al.
Pubblicazione: (2026)
Memory DisOrder: Memory Re-orderings as a Timerless Side-channel
di: Siddens, Sean, et al.
Pubblicazione: (2026)
di: Siddens, Sean, et al.
Pubblicazione: (2026)
LP-GEMM: Integrating Layout Propagation into GEMM Operations
di: Carneiro, César Guedes, et al.
Pubblicazione: (2026)
di: Carneiro, César Guedes, et al.
Pubblicazione: (2026)
Harnessing Batched BLAS/LAPACK Kernels on GPUs for Parallel Solutions of Block Tridiagonal Systems
di: Jin, David, et al.
Pubblicazione: (2025)
di: Jin, David, et al.
Pubblicazione: (2025)
AutoTriton: Automatic Triton Programming with Reinforcement Learning in LLMs
di: Li, Shangzhan, et al.
Pubblicazione: (2025)
di: Li, Shangzhan, et al.
Pubblicazione: (2025)
Seawater carbonate chemistry and dissolution of the triton shell
di: Harvey, Ben P, et al.
Pubblicazione: (2018)
di: Harvey, Ben P, et al.
Pubblicazione: (2018)
DRTriton: Large-Scale Synthetic Data Driven Reinforcement Learning for Triton Kernel Generation
di: Guo, Siqi, et al.
Pubblicazione: (2026)
di: Guo, Siqi, et al.
Pubblicazione: (2026)
AutoKernel: Autonomous GPU Kernel Optimization via Iterative Agent-Driven Search
di: Jaber, Jaber, et al.
Pubblicazione: (2026)
di: Jaber, Jaber, et al.
Pubblicazione: (2026)
TritonRL: Training LLMs to Think and Code Triton Without Cheating
di: Woo, Jiin, et al.
Pubblicazione: (2025)
di: Woo, Jiin, et al.
Pubblicazione: (2025)
EmuGEMM (Benchmark Scripts): Fused Tensor Core Kernels for Precision Emulation in Matrix Multiplication
di: Lu, Denghui, et al.
Pubblicazione: (2026)
di: Lu, Denghui, et al.
Pubblicazione: (2026)
Matrix-Free 3D SIMP Topology Optimization with Fused Gather-GEMM-Scatter Kernels
di: Yang, Shaoliang, et al.
Pubblicazione: (2026)
di: Yang, Shaoliang, et al.
Pubblicazione: (2026)
Scaling Analog Photonic Accelerators for Byte-Size, Integer General Matrix Multiply (GEMM) Kernels
di: Alo, Oluwaseun Adewunmi, et al.
Pubblicazione: (2024)
di: Alo, Oluwaseun Adewunmi, et al.
Pubblicazione: (2024)
CodeGEMM: A Codebook-Centric Approach to Efficient GEMM in Quantized LLMs
di: Park, Gunho, et al.
Pubblicazione: (2025)
di: Park, Gunho, et al.
Pubblicazione: (2025)
Determining Pasteur Parameter for Chiral Medium using Long‐Range Surface Plasmon Resonance
di: Lixun Sun, et al.
Pubblicazione: (2024)
di: Lixun Sun, et al.
Pubblicazione: (2024)
Hate Speech Law
di: Brown, Alex
Pubblicazione: (2025)
di: Brown, Alex
Pubblicazione: (2025)
Kernel splitting, Streams, cuBLAS (work in progress in the MG5aMC CUDACPP plugin)
di: Valassi, Andrea
Pubblicazione: (2024)
di: Valassi, Andrea
Pubblicazione: (2024)
Neutron-neutron distribution of the triton from pionless EFT
di: Kirchner, Tanja, et al.
Pubblicazione: (2024)
di: Kirchner, Tanja, et al.
Pubblicazione: (2024)
The Humanities at Triton College.
di: Jacot, Robert E., et al.
Pubblicazione: (1984)
di: Jacot, Robert E., et al.
Pubblicazione: (1984)
GraphBLAS Mathematical Opportunities: Parallel Hypersparse, Matrix Based Graph Streaming, and Complex-Index Matrices
di: Jananthan, Hayden, et al.
Pubblicazione: (2025)
di: Jananthan, Hayden, et al.
Pubblicazione: (2025)
TritonBench: Benchmarking Large Language Model Capabilities for Generating Triton Operators
di: Li, Jianling, et al.
Pubblicazione: (2025)
di: Li, Jianling, et al.
Pubblicazione: (2025)
Analytical Proofs for JSD Contraction Coefficients of AWGN Channels
di: Alex Shvets
Pubblicazione: (2025)
di: Alex Shvets
Pubblicazione: (2025)
Minimization Principle for Analytical Solution of Turbulent Flow in Channel
di: Fedoseyev, Alex
Pubblicazione: (2024)
di: Fedoseyev, Alex
Pubblicazione: (2024)
ML-Triton, A Multi-Level Compilation and Language Extension to Triton GPU Programming
di: Wang, Dewei, et al.
Pubblicazione: (2025)
di: Wang, Dewei, et al.
Pubblicazione: (2025)
GEMM-GS: Accelerating 3D Gaussian Splatting on Tensor Cores with GEMM-Compatible Blending
di: Li, Haomin, et al.
Pubblicazione: (2026)
di: Li, Haomin, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Fleet: Hierarchical Task-based Abstraction for Megakernels on Multi-Die GPUs
di: Chowdhary, Sangeeta, et al.
Pubblicazione: (2026) -
HipKittens: Fast and Furious AMD Kernels
di: Hu, William, et al.
Pubblicazione: (2025) -
SwizzlePerf: Hardware-Aware LLMs for GPU Kernel Performance Optimization
di: Tschand, Arya, et al.
Pubblicazione: (2025) -
Seer: Predictive Runtime Kernel Selection for Irregular Problems
di: Swann, Ryan, et al.
Pubblicazione: (2024) -
TritonForge: Profiling-Guided Framework for Automated Triton Kernel Optimization
di: Li, Haonan, et al.
Pubblicazione: (2025)