KForge: Program Synthesis for Diverse AI Hardware Accelerators

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sereda, Taras, John, Tom St., Bartan, Burak, Serrino, Natalie, Katti, Sachin, Asgar, Zain
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908659419709440
author Sereda, Taras
John, Tom St.
Bartan, Burak
Serrino, Natalie
Katti, Sachin
Asgar, Zain
author_facet Sereda, Taras
John, Tom St.
Bartan, Burak
Serrino, Natalie
Katti, Sachin
Asgar, Zain
contents GPU kernels are critical for ML performance but difficult to optimize across diverse accelerators. We present KForge, a platform-agnostic framework built on two collaborative LLM-based agents: a generation agent that produces and iteratively refines programs through compilation and correctness feedback, and a performance analysis agent that interprets profiling data to guide optimization. This agent-based architecture requires only a single-shot example to target new platforms. We make three key contributions: (1) introducing an iterative refinement system where the generation agent and performance analysis agent collaborate through functional and optimization passes, interpreting diverse profiling data (from programmatic APIs to GUI-based tools) to generate actionable recommendations that guide program synthesis for arbitrary accelerators; (2) demonstrating that the generation agent effectively leverages cross-platform knowledge transfer, where a reference implementation from one architecture substantially improves generation quality for different hardware targets; and (3) validating the platform-agnostic nature of our approach by demonstrating effective program synthesis across fundamentally different parallel computing platforms: NVIDIA CUDA and Apple Metal.
format Preprint
id arxiv_https___arxiv_org_abs_2511_13274
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle KForge: Program Synthesis for Diverse AI Hardware Accelerators
Sereda, Taras
John, Tom St.
Bartan, Burak
Serrino, Natalie
Katti, Sachin
Asgar, Zain
Machine Learning
Artificial Intelligence
Multiagent Systems
Performance
Software Engineering
GPU kernels are critical for ML performance but difficult to optimize across diverse accelerators. We present KForge, a platform-agnostic framework built on two collaborative LLM-based agents: a generation agent that produces and iteratively refines programs through compilation and correctness feedback, and a performance analysis agent that interprets profiling data to guide optimization. This agent-based architecture requires only a single-shot example to target new platforms. We make three key contributions: (1) introducing an iterative refinement system where the generation agent and performance analysis agent collaborate through functional and optimization passes, interpreting diverse profiling data (from programmatic APIs to GUI-based tools) to generate actionable recommendations that guide program synthesis for arbitrary accelerators; (2) demonstrating that the generation agent effectively leverages cross-platform knowledge transfer, where a reference implementation from one architecture substantially improves generation quality for different hardware targets; and (3) validating the platform-agnostic nature of our approach by demonstrating effective program synthesis across fundamentally different parallel computing platforms: NVIDIA CUDA and Apple Metal.
title KForge: Program Synthesis for Diverse AI Hardware Accelerators
topic Machine Learning
Artificial Intelligence
Multiagent Systems
Performance
Software Engineering
url https://arxiv.org/abs/2511.13274