A Performance Model for Warp Specialization Kernels
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Zhengyang, Grover, Vinod |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Tawa: Automatic Warp Specialization for Modern GPUs with Asynchronous References
von: Chen, Hongzheng, et al.
Veröffentlicht: (2025)
von: Chen, Hongzheng, et al.
Veröffentlicht: (2025)
Modeling Layout Abstractions Using Integer Set Relations
von: Bhaskaracharya, Somashekaracharya G, et al.
Veröffentlicht: (2025)
von: Bhaskaracharya, Somashekaracharya G, et al.
Veröffentlicht: (2025)
Optimal Software Pipelining and Warp Specialization for Tensor Core GPUs
von: Soi, Rupanshu, et al.
Veröffentlicht: (2025)
von: Soi, Rupanshu, et al.
Veröffentlicht: (2025)
Pattern Matching in AI Compilers and its Formalization (Extended Version)
von: Cutler, Joseph W., et al.
Veröffentlicht: (2024)
von: Cutler, Joseph W., et al.
Veröffentlicht: (2024)
Minotaur: A SIMD-Oriented Synthesizing Superoptimizer
von: Liu, Zhengyang, et al.
Veröffentlicht: (2023)
von: Liu, Zhengyang, et al.
Veröffentlicht: (2023)
Scaling Deep Learning Training with MPMD Pipeline Parallelism
von: Xhebraj, Anxhelo, et al.
Veröffentlicht: (2024)
von: Xhebraj, Anxhelo, et al.
Veröffentlicht: (2024)
Model2Kernel: Model-Aware Symbolic Execution For Safe CUDA Kernels
von: He, Mengting, et al.
Veröffentlicht: (2026)
von: He, Mengting, et al.
Veröffentlicht: (2026)
Equivalence Checking of ML GPU Kernels
von: Dubey, Kshitij, et al.
Veröffentlicht: (2025)
von: Dubey, Kshitij, et al.
Veröffentlicht: (2025)
Kernel Contracts: A Specification Language for ML Kernel Correctness Across Heterogeneous Silicon
von: Veit, Cooper
Veröffentlicht: (2026)
von: Veit, Cooper
Veröffentlicht: (2026)
Synthesizing Formal Semantics from Executable Interpreters
von: Liu, Jiangyi, et al.
Veröffentlicht: (2024)
von: Liu, Jiangyi, et al.
Veröffentlicht: (2024)
NPUEval: Optimizing NPU Kernels with LLMs and Open Source Compilers
von: Kalade, Sarunas, et al.
Veröffentlicht: (2025)
von: Kalade, Sarunas, et al.
Veröffentlicht: (2025)
Special Delivery: Programming with Mailbox Types (Extended Version)
von: Fowler, Simon, et al.
Veröffentlicht: (2023)
von: Fowler, Simon, et al.
Veröffentlicht: (2023)
Flex Attention: A Programming Model for Generating Optimized Attention Kernels
von: Dong, Juechu, et al.
Veröffentlicht: (2024)
von: Dong, Juechu, et al.
Veröffentlicht: (2024)
Insum: Sparse GPU Kernels Simplified and Optimized with Indirect Einsums
von: Won, Jaeyeon, et al.
Veröffentlicht: (2025)
von: Won, Jaeyeon, et al.
Veröffentlicht: (2025)
KPerfIR: Towards an Open and Compiler-centric Ecosystem for GPU Kernel Performance Tooling on Modern AI Workloads
von: Guan, Yue, et al.
Veröffentlicht: (2025)
von: Guan, Yue, et al.
Veröffentlicht: (2025)
Kernel-FFI: Transparent Foreign Function Interfaces for Interactive Notebooks
von: Li, Hebi, et al.
Veröffentlicht: (2025)
von: Li, Hebi, et al.
Veröffentlicht: (2025)
KAIJU: An Executive Kernel for Intent-Gated Execution of LLM Agents
von: Guerin, Cormac, et al.
Veröffentlicht: (2026)
von: Guerin, Cormac, et al.
Veröffentlicht: (2026)
Mirage Persistent Kernel: A Compiler and Runtime for Mega-Kernelizing Tensor Programs
von: Cheng, Xinhao, et al.
Veröffentlicht: (2025)
von: Cheng, Xinhao, et al.
Veröffentlicht: (2025)
Nautilus: An Auto-Scheduling Tensor Compiler for Efficient Tiled GPU Kernels
von: Zhao, Yifan, et al.
Veröffentlicht: (2026)
von: Zhao, Yifan, et al.
Veröffentlicht: (2026)
Analyzing Latency Hiding and Parallelism in an MLIR-based AI Kernel Compiler
von: Absar, Javed, et al.
Veröffentlicht: (2026)
von: Absar, Javed, et al.
Veröffentlicht: (2026)
Code Generation for Cryptographic Kernels using Multi-word Modular Arithmetic on GPU
von: Zhang, Naifeng, et al.
Veröffentlicht: (2025)
von: Zhang, Naifeng, et al.
Veröffentlicht: (2025)
Correctness is Demanding, Performance is Frustrating
von: Sinkarovs, Artjoms, et al.
Veröffentlicht: (2024)
von: Sinkarovs, Artjoms, et al.
Veröffentlicht: (2024)
BODHI: Precise OS Kernel Specification Inference
von: Chang, Zhiming, et al.
Veröffentlicht: (2026)
von: Chang, Zhiming, et al.
Veröffentlicht: (2026)
Towards a Scalable Proof Engine: A Performant Prototype Rewriting Primitive for Coq
von: Gross, Jason, et al.
Veröffentlicht: (2023)
von: Gross, Jason, et al.
Veröffentlicht: (2023)
Quantifying the Importance of Data Alignment in Downstream Model Performance
von: Chawla, Krrish, et al.
Veröffentlicht: (2025)
von: Chawla, Krrish, et al.
Veröffentlicht: (2025)
AutoLALA: Automatic Loop Algebraic Locality Analysis for AI and HPC Kernels
von: Zhu, Yifan, et al.
Veröffentlicht: (2026)
von: Zhu, Yifan, et al.
Veröffentlicht: (2026)
Very High Level Programming Languages (e.g., SNOBOL, COMIT) in the Special Librarian's Future
von: Libbey, Miles A.
Veröffentlicht: (1975)
von: Libbey, Miles A.
Veröffentlicht: (1975)
Fail Faster: Staging and Fast Randomness for High-Performance PBT
von: Richey, Cynthia, et al.
Veröffentlicht: (2025)
von: Richey, Cynthia, et al.
Veröffentlicht: (2025)
Efficient Selection of Type Annotations for Performance Improvement in Gradual Typing
von: Li, Senxi, et al.
Veröffentlicht: (2026)
von: Li, Senxi, et al.
Veröffentlicht: (2026)
QPanda3: A High-Performance Software-Hardware Collaborative Framework for Large-Scale Quantum-Classical Computing Integration
von: Zou, Tianrui, et al.
Veröffentlicht: (2025)
von: Zou, Tianrui, et al.
Veröffentlicht: (2025)
OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models
von: Huang, Siming, et al.
Veröffentlicht: (2024)
von: Huang, Siming, et al.
Veröffentlicht: (2024)
Dr Wenowdis: Specializing dynamic language C extensions using type information
von: Bernstein, Maxwell, et al.
Veröffentlicht: (2024)
von: Bernstein, Maxwell, et al.
Veröffentlicht: (2024)
Stencil-Lifting: Hierarchical Recursive Lifting System for Extracting Summary of Stencil Kernel in Legacy Codes
von: Li, Mingyi, et al.
Veröffentlicht: (2025)
von: Li, Mingyi, et al.
Veröffentlicht: (2025)
Owi: Performant Parallel Symbolic Execution Made Easy, an Application to WebAssembly
von: Andrès, Léo, et al.
Veröffentlicht: (2024)
von: Andrès, Léo, et al.
Veröffentlicht: (2024)
SUQL: Conversational Search over Structured and Unstructured Data with Large Language Models
von: Liu, Shicheng, et al.
Veröffentlicht: (2023)
von: Liu, Shicheng, et al.
Veröffentlicht: (2023)
Extended Abstract: Towards a Performance Comparison of Syntax and Type-Directed NbE
von: Gould, Chester J. F., et al.
Veröffentlicht: (2025)
von: Gould, Chester J. F., et al.
Veröffentlicht: (2025)
Fully Symbolic Analysis of Loop Locality: Using Imaginary Reuse to Infer Real Performance
von: Zhu, Yifan, et al.
Veröffentlicht: (2026)
von: Zhu, Yifan, et al.
Veröffentlicht: (2026)
Data-efficient Performance Modeling via Pre-training
von: Liu, Chunting, et al.
Veröffentlicht: (2025)
von: Liu, Chunting, et al.
Veröffentlicht: (2025)
A Brief Survey of Formal Models of Concurrency
von: Averill, Charles
Veröffentlicht: (2024)
von: Averill, Charles
Veröffentlicht: (2024)
ALEA IACTA EST: A Declarative Domain-Specific Language for Manually Performable Random Experiments
von: Widemann, Baltasar Trancón y, et al.
Veröffentlicht: (2025)
von: Widemann, Baltasar Trancón y, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Tawa: Automatic Warp Specialization for Modern GPUs with Asynchronous References
von: Chen, Hongzheng, et al.
Veröffentlicht: (2025) -
Modeling Layout Abstractions Using Integer Set Relations
von: Bhaskaracharya, Somashekaracharya G, et al.
Veröffentlicht: (2025) -
Optimal Software Pipelining and Warp Specialization for Tensor Core GPUs
von: Soi, Rupanshu, et al.
Veröffentlicht: (2025) -
Pattern Matching in AI Compilers and its Formalization (Extended Version)
von: Cutler, Joseph W., et al.
Veröffentlicht: (2024) -
Minotaur: A SIMD-Oriented Synthesizing Superoptimizer
von: Liu, Zhengyang, et al.
Veröffentlicht: (2023)