Record-Remix-Replay: Hierarchical GPU Kernel Optimization using Evolutionary Search
Fuente:
arXiv
Salvato in:
| Autori principali: | Nichols, Daniel, Parasyris, Konstantinos, Melone, Caetano, Ben-Nun, Tal, Georgakoudis, Giorgis, Menon, Harshitha |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Integrating Performance Tools in Model Reasoning for GPU Kernel Optimization
di: Nichols, Daniel, et al.
Pubblicazione: (2025)
di: Nichols, Daniel, et al.
Pubblicazione: (2025)
Taking GPU Programming Models to Task for Performance Portability
di: Davis, Joshua H., et al.
Pubblicazione: (2024)
di: Davis, Joshua H., et al.
Pubblicazione: (2024)
Counting Without Running: Evaluating LLMs' Reasoning About Code Complexity
di: Bolet, Gregory, et al.
Pubblicazione: (2025)
di: Bolet, Gregory, et al.
Pubblicazione: (2025)
Can Large Language Models Predict Parallel Code Performance?
di: Bolet, Gregory, et al.
Pubblicazione: (2025)
di: Bolet, Gregory, et al.
Pubblicazione: (2025)
LLMs as Packagers of HPC Software
di: Melone, Caetano, et al.
Pubblicazione: (2025)
di: Melone, Caetano, et al.
Pubblicazione: (2025)
HPAC-ML: A Programming Model for Embedding ML Surrogates in Scientific Applications
di: Fink, Zane, et al.
Pubblicazione: (2024)
di: Fink, Zane, et al.
Pubblicazione: (2024)
Inductive Loop Analysis for Practical HPC Application Optimization
di: Schaad, Philipp, et al.
Pubblicazione: (2025)
di: Schaad, Philipp, et al.
Pubblicazione: (2025)
KEET: Explaining Performance of GPU Kernels Using LLM Agents
di: Davis, Joshua H., et al.
Pubblicazione: (2026)
di: Davis, Joshua H., et al.
Pubblicazione: (2026)
Leveraging AI for Productive and Trustworthy HPC Software: Challenges and Research Directions
di: Teranishi, Keita, et al.
Pubblicazione: (2025)
di: Teranishi, Keita, et al.
Pubblicazione: (2025)
Optimizing Agentic Language Model Inference via Speculative Tool Calls
di: Nichols, Daniel, et al.
Pubblicazione: (2025)
di: Nichols, Daniel, et al.
Pubblicazione: (2025)
Data-Driven Analysis to Understand GPU Hardware Resource Usage of Optimizations
di: Islam, Tanzima Z., et al.
Pubblicazione: (2024)
di: Islam, Tanzima Z., et al.
Pubblicazione: (2024)
The Energy Cost of Execution-Idle in GPU Clusters
di: Lei, Yiran, et al.
Pubblicazione: (2026)
di: Lei, Yiran, et al.
Pubblicazione: (2026)
Scalable GPU Performance Variability Analysis framework
di: Lahiry, Ankur, et al.
Pubblicazione: (2025)
di: Lahiry, Ankur, et al.
Pubblicazione: (2025)
On the Partitioning of GPU Power among Multi-Instances
di: Vamja, Tirth, et al.
Pubblicazione: (2025)
di: Vamja, Tirth, et al.
Pubblicazione: (2025)
Disaggregated Design for GPU-Based Volumetric Data Structures
di: Meneghin, Massimiliano, et al.
Pubblicazione: (2025)
di: Meneghin, Massimiliano, et al.
Pubblicazione: (2025)
Arkade: k-Nearest Neighbor Search With Non-Euclidean Distances using GPU Ray Tracing
di: Mandarapu, Durga, et al.
Pubblicazione: (2023)
di: Mandarapu, Durga, et al.
Pubblicazione: (2023)
Profiling and optimization of multi-card GPU machine learning jobs
di: Lawenda, Marcin, et al.
Pubblicazione: (2025)
di: Lawenda, Marcin, et al.
Pubblicazione: (2025)
CUTHERMO: Understanding GPU Memory Inefficiencies with Heat Map Profiling
di: Zhao, Yanbo, et al.
Pubblicazione: (2025)
di: Zhao, Yanbo, et al.
Pubblicazione: (2025)
High-Performance Portable GPU Primitives for Arbitrary Types and Operators in Julia
di: Pilliat, Emmanuel
Pubblicazione: (2026)
di: Pilliat, Emmanuel
Pubblicazione: (2026)
Efficient allocation of image recognition and LLM tasks on multi-GPU system
di: Lawenda, Marcin, et al.
Pubblicazione: (2025)
di: Lawenda, Marcin, et al.
Pubblicazione: (2025)
Efficient GPU-Centered Singular Value Decomposition Using the Divide-and-Conquer Method
di: Liu, Shifang, et al.
Pubblicazione: (2025)
di: Liu, Shifang, et al.
Pubblicazione: (2025)
Minos: Systematically Classifying Performance and Power Characteristics of GPU Workloads on HPC Clusters
di: Jain, Rutwik, et al.
Pubblicazione: (2026)
di: Jain, Rutwik, et al.
Pubblicazione: (2026)
Unleashing the Power of Preemptive Priority-based Scheduling for Real-Time GPU Tasks
di: Wang, Yidi, et al.
Pubblicazione: (2024)
di: Wang, Yidi, et al.
Pubblicazione: (2024)
Recorder: Comprehensive Parallel I/O Tracing and Analysis
di: Wang, Chen, et al.
Pubblicazione: (2025)
di: Wang, Chen, et al.
Pubblicazione: (2025)
Can Tensor Cores Benefit Memory-Bound Kernels? (No!)
di: Zhang, Lingqi, et al.
Pubblicazione: (2025)
di: Zhang, Lingqi, et al.
Pubblicazione: (2025)
Dissecting CPU-GPU Unified Physical Memory on AMD MI300A APUs
di: Wahlgren, Jacob, et al.
Pubblicazione: (2025)
di: Wahlgren, Jacob, et al.
Pubblicazione: (2025)
LEO: Tracing GPU Stall Root Causes via Cross-Vendor Backward Slicing
di: Xia, Yuning, et al.
Pubblicazione: (2026)
di: Xia, Yuning, et al.
Pubblicazione: (2026)
Fast and Scalable Mixed Precision Euclidean Distance Calculations Using GPU Tensor Cores
di: Curless, Brian, et al.
Pubblicazione: (2025)
di: Curless, Brian, et al.
Pubblicazione: (2025)
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
di: Lin, Mao, et al.
Pubblicazione: (2026)
di: Lin, Mao, et al.
Pubblicazione: (2026)
Cache Blocking of Distributed-Memory Parallel Matrix Power Kernels
di: Lacey, Dane C., et al.
Pubblicazione: (2024)
di: Lacey, Dane C., et al.
Pubblicazione: (2024)
FACT: Compositional Kernel Synthesis with a Three-Stage Agentic Workflow
di: Heidari, Sina, et al.
Pubblicazione: (2026)
di: Heidari, Sina, et al.
Pubblicazione: (2026)
Towards Portability at Scale: A Cross-Architecture Performance Evaluation of a GPU-enabled Shallow Water Solver
di: Villalobos, Johansell, et al.
Pubblicazione: (2025)
di: Villalobos, Johansell, et al.
Pubblicazione: (2025)
Automated Programmatic Performance Analysis of Parallel Programs
di: Cankur, Onur, et al.
Pubblicazione: (2024)
di: Cankur, Onur, et al.
Pubblicazione: (2024)
AcceleratedKernels.jl: Cross-Architecture Parallel Algorithms from a Unified, Transpiled Codebase
di: Nicusan, Andrei-Leonard, et al.
Pubblicazione: (2025)
di: Nicusan, Andrei-Leonard, et al.
Pubblicazione: (2025)
Portable High-Performance Kernel Generation for a Computational Fluid Dynamics Code with DaCe
di: Andersson, Måns I., et al.
Pubblicazione: (2025)
di: Andersson, Måns I., et al.
Pubblicazione: (2025)
GPU Kernel Optimization Beyond Full Builds: An LLM Framework with Minimal Executable Programs
di: Chu, Ruifan, et al.
Pubblicazione: (2025)
di: Chu, Ruifan, et al.
Pubblicazione: (2025)
Cloud Resource Allocation with Convex Optimization
di: Boghani, Shayan, et al.
Pubblicazione: (2025)
di: Boghani, Shayan, et al.
Pubblicazione: (2025)
Bringing Auto-tuning to HIP: Analysis of Tuning Impact and Difficulty on AMD and Nvidia GPUs
di: Lurati, Milo, et al.
Pubblicazione: (2024)
di: Lurati, Milo, et al.
Pubblicazione: (2024)
Optimizing Near Field Computation in the MLFMA Algorithm with Data Redundancy and Performance Modeling on a Single GPU
di: Sadeghi, Morteza, et al.
Pubblicazione: (2024)
di: Sadeghi, Morteza, et al.
Pubblicazione: (2024)
The Landscape of GPU-Centric Communication
di: Unat, Didem, et al.
Pubblicazione: (2024)
di: Unat, Didem, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Integrating Performance Tools in Model Reasoning for GPU Kernel Optimization
di: Nichols, Daniel, et al.
Pubblicazione: (2025) -
Taking GPU Programming Models to Task for Performance Portability
di: Davis, Joshua H., et al.
Pubblicazione: (2024) -
Counting Without Running: Evaluating LLMs' Reasoning About Code Complexity
di: Bolet, Gregory, et al.
Pubblicazione: (2025) -
Can Large Language Models Predict Parallel Code Performance?
di: Bolet, Gregory, et al.
Pubblicazione: (2025) -
LLMs as Packagers of HPC Software
di: Melone, Caetano, et al.
Pubblicazione: (2025)