VeriLocc: End-to-End Cross-Architecture Register Allocation via LLM

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Jin, Lesheng, Ruan, Zhenyuan, Mai, Haohui, Shang, Jingbo
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913906487721984
author Jin, Lesheng
Ruan, Zhenyuan
Mai, Haohui
Shang, Jingbo
author_facet Jin, Lesheng
Ruan, Zhenyuan
Mai, Haohui
Shang, Jingbo
contents Modern GPUs evolve rapidly, yet production compilers still rely on hand-crafted register allocation heuristics that require substantial re-tuning for each hardware generation. We introduce VeriLocc, a framework that combines large language models (LLMs) with formal compiler techniques to enable generalizable and verifiable register allocation across GPU architectures. VeriLocc fine-tunes an LLM to translate intermediate representations (MIRs) into target-specific register assignments, aided by static analysis for cross-architecture normalization and generalization and a verifier-guided regeneration loop to ensure correctness. Evaluated on matrix multiplication (GEMM) and multi-head attention (MHA), VeriLocc achieves 85-99% single-shot accuracy and near-100% pass@100. Case study shows that VeriLocc discovers more performant assignments than expert-tuned libraries, outperforming rocBLAS by over 10% in runtime.
format Preprint
id arxiv_https___arxiv_org_abs_2506_17506
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VeriLocc: End-to-End Cross-Architecture Register Allocation via LLM
Jin, Lesheng
Ruan, Zhenyuan
Mai, Haohui
Shang, Jingbo
Computation and Language
Operating Systems
Modern GPUs evolve rapidly, yet production compilers still rely on hand-crafted register allocation heuristics that require substantial re-tuning for each hardware generation. We introduce VeriLocc, a framework that combines large language models (LLMs) with formal compiler techniques to enable generalizable and verifiable register allocation across GPU architectures. VeriLocc fine-tunes an LLM to translate intermediate representations (MIRs) into target-specific register assignments, aided by static analysis for cross-architecture normalization and generalization and a verifier-guided regeneration loop to ensure correctness. Evaluated on matrix multiplication (GEMM) and multi-head attention (MHA), VeriLocc achieves 85-99% single-shot accuracy and near-100% pass@100. Case study shows that VeriLocc discovers more performant assignments than expert-tuned libraries, outperforming rocBLAS by over 10% in runtime.
title VeriLocc: End-to-End Cross-Architecture Register Allocation via LLM
topic Computation and Language
Operating Systems
url https://arxiv.org/abs/2506.17506