HPCTransCompile: An AI Compiler Generated Dataset for High-Performance CUDA Transpilation and LLM Preliminary Exploration
Fuente:
arXiv
Salvato in:
| Autori principali: | Lv, Jiaqi, He, Xufeng, Liu, Yanchen, Dai, Xu, Shen, Aocheng, Li, Yinghao, Hao, Jiachen, Ding, Jianrong, Hu, Yang, Yin, Shouyi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DeepCompile: A Compiler-Driven Approach to Optimizing Distributed Deep Learning Training
di: Tanaka, Masahiro, et al.
Pubblicazione: (2025)
di: Tanaka, Masahiro, et al.
Pubblicazione: (2025)
Flint: Compiler Enabled Cluster-Free Design Space Exploration for Distributed ML
di: Yoo, Jinsun, et al.
Pubblicazione: (2026)
di: Yoo, Jinsun, et al.
Pubblicazione: (2026)
Efficient Parallel Compilation and Profiling of Quantum Circuits at Large Scales
di: Moore, Jane, et al.
Pubblicazione: (2026)
di: Moore, Jane, et al.
Pubblicazione: (2026)
LAPIS: A Performance Portable, High Productivity Compiler Framework
di: Kelley, Brian, et al.
Pubblicazione: (2025)
di: Kelley, Brian, et al.
Pubblicazione: (2025)
Zen-Attention: A Compiler Framework for Dynamic Attention Folding on AMD NPUs
di: Deshmukh, Aadesh, et al.
Pubblicazione: (2025)
di: Deshmukh, Aadesh, et al.
Pubblicazione: (2025)
Gensor: A Graph-based Construction Tensor Compilation Method for Deep Learning
di: Liu, Hangda, et al.
Pubblicazione: (2025)
di: Liu, Hangda, et al.
Pubblicazione: (2025)
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler
di: Zheng, Size, et al.
Pubblicazione: (2025)
di: Zheng, Size, et al.
Pubblicazione: (2025)
nncase: An End-to-End Compiler for Efficient LLM Deployment on Heterogeneous Storage Architectures
di: Guo, Hui, et al.
Pubblicazione: (2025)
di: Guo, Hui, et al.
Pubblicazione: (2025)
Parallel Gaussian process with kernel approximation in CUDA
di: Carminati, Davide
Pubblicazione: (2024)
di: Carminati, Davide
Pubblicazione: (2024)
Tutoring LLM into a Better CUDA Optimizer
di: Brabec, Matyáš, et al.
Pubblicazione: (2025)
di: Brabec, Matyáš, et al.
Pubblicazione: (2025)
Making Serverless Computing Extensible: A Case Study of Serverless Data Analytics
di: Yu, Minchen, et al.
Pubblicazione: (2025)
di: Yu, Minchen, et al.
Pubblicazione: (2025)
Comparison of Vectorization Capabilities of Different Compilers for X86 and ARM CPUs
di: Sakib, Nazmus, et al.
Pubblicazione: (2025)
di: Sakib, Nazmus, et al.
Pubblicazione: (2025)
PALM: A Efficient Performance Simulator for Tiled Accelerators with Large-scale Model Training
di: Fang, Jiahao, et al.
Pubblicazione: (2024)
di: Fang, Jiahao, et al.
Pubblicazione: (2024)
High-Performance Parallelization of Dijkstra's Algorithm Using MPI and CUDA
di: Song, Boyang
Pubblicazione: (2025)
di: Song, Boyang
Pubblicazione: (2025)
VTC: DNN Compilation with Virtual Tensors for Data Movement Elimination
di: Hu, Muyan, et al.
Pubblicazione: (2026)
di: Hu, Muyan, et al.
Pubblicazione: (2026)
Parallel DNA Sequence Alignment on High-Performance Systems with CUDA and MPI
di: Zwaka, Linus
Pubblicazione: (2024)
di: Zwaka, Linus
Pubblicazione: (2024)
Boosting Performance of Iterative Applications on GPUs: Kernel Batching with CUDA Graphs
di: Ekelund, Jonah, et al.
Pubblicazione: (2025)
di: Ekelund, Jonah, et al.
Pubblicazione: (2025)
ZipFlow: a Compiler-based Framework to Unleash Compressed Data Movement for Modern GPUs
di: Yeo, Gwangoo, et al.
Pubblicazione: (2026)
di: Yeo, Gwangoo, et al.
Pubblicazione: (2026)
ECDQC: Efficient Compilation for Distributed Quantum Computing with Linear Layout
di: Liu, Kecheng, et al.
Pubblicazione: (2024)
di: Liu, Kecheng, et al.
Pubblicazione: (2024)
Optimizing Compilation for Distributed Quantum Computing via Clustering and Annealing
di: Zhou, Ruilin, et al.
Pubblicazione: (2025)
di: Zhou, Ruilin, et al.
Pubblicazione: (2025)
Efficient Gate Reordering for Distributed Quantum Compiling in Data Centers
di: Mengoni, Riccardo, et al.
Pubblicazione: (2025)
di: Mengoni, Riccardo, et al.
Pubblicazione: (2025)
Parallel Paradigms in Modern HPC: A Comparative Analysis of MPI, OpenMP, and CUDA
di: ALHafez, Nizar, et al.
Pubblicazione: (2025)
di: ALHafez, Nizar, et al.
Pubblicazione: (2025)
Atomique: A Quantum Compiler for Reconfigurable Neutral Atom Arrays
di: Wang, Hanrui, et al.
Pubblicazione: (2023)
di: Wang, Hanrui, et al.
Pubblicazione: (2023)
KPerfIR: Towards an Open and Compiler-centric Ecosystem for GPU Kernel Performance Tooling on Modern AI Workloads
di: Guan, Yue, et al.
Pubblicazione: (2025)
di: Guan, Yue, et al.
Pubblicazione: (2025)
Lessons Learned Migrating CUDA to SYCL: A HEP Case Study with ROOT RDataFrame
di: Chen, Jolly, et al.
Pubblicazione: (2024)
di: Chen, Jolly, et al.
Pubblicazione: (2024)
Accelerating Intra-Node GPU-to-GPU Communication Through Multi-Path Transfers with CUDA Graphs
di: Sojoodi, Amirhossein, et al.
Pubblicazione: (2026)
di: Sojoodi, Amirhossein, et al.
Pubblicazione: (2026)
AutoSP: Unlocking Long-Context LLM Training Via Compiler-Based Sequence Parallelism
di: Gupta, Ahan, et al.
Pubblicazione: (2026)
di: Gupta, Ahan, et al.
Pubblicazione: (2026)
AcceleratedKernels.jl: Cross-Architecture Parallel Algorithms from a Unified, Transpiled Codebase
di: Nicusan, Andrei-Leonard, et al.
Pubblicazione: (2025)
di: Nicusan, Andrei-Leonard, et al.
Pubblicazione: (2025)
Leveraging Neural Graph Compilers in Machine Learning Research for Edge-Cloud Systems
di: Furutanpey, Alireza, et al.
Pubblicazione: (2025)
di: Furutanpey, Alireza, et al.
Pubblicazione: (2025)
Compiler Support for Speculation in Decoupled Access/Execute Architectures
di: Szafarczyk, Robert, et al.
Pubblicazione: (2025)
di: Szafarczyk, Robert, et al.
Pubblicazione: (2025)
ScanWeaver: Compiler-Driven Parallelization of Affine Recurrences via Associative Scan Lowering
di: Wu, Qiying, et al.
Pubblicazione: (2026)
di: Wu, Qiying, et al.
Pubblicazione: (2026)
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start
di: Liu, Xueshen, et al.
Pubblicazione: (2026)
di: Liu, Xueshen, et al.
Pubblicazione: (2026)
cuVegas: Accelerate Multidimensional Monte Carlo Integration through a Parallelized CUDA-based Implementation of the VEGAS Enhanced Algorithm
di: Tolotti, Emiliano, et al.
Pubblicazione: (2024)
di: Tolotti, Emiliano, et al.
Pubblicazione: (2024)
Resource-Efficient Compilation of Distributed Quantum Circuits for Solving Large-Scale Wireless Communication Network Problems
di: Chen, Kuan-Cheng, et al.
Pubblicazione: (2025)
di: Chen, Kuan-Cheng, et al.
Pubblicazione: (2025)
Deep Learning and Machine Learning with GPGPU and CUDA: Unlocking the Power of Parallel Computing
di: Li, Ming, et al.
Pubblicazione: (2024)
di: Li, Ming, et al.
Pubblicazione: (2024)
cuConv: A CUDA Implementation of Convolution for CNN Inference
di: Jordà, Marc, et al.
Pubblicazione: (2021)
di: Jordà, Marc, et al.
Pubblicazione: (2021)
A Comparison of the Performance of the Molecular Dynamics Simulation Package GROMACS Implemented in the SYCL and CUDA Programming Models
di: Apanasevich, L., et al.
Pubblicazione: (2024)
di: Apanasevich, L., et al.
Pubblicazione: (2024)
LOOPer: A Learned Automatic Code Optimizer For Polyhedral Compilers
di: Merouani, Massinissa, et al.
Pubblicazione: (2024)
di: Merouani, Massinissa, et al.
Pubblicazione: (2024)
Theoretical Foundations of GPU-Native Compilation for Rapid Code Iteration
di: Metinov, Adilet, et al.
Pubblicazione: (2025)
di: Metinov, Adilet, et al.
Pubblicazione: (2025)
Event Tensor: A Unified Abstraction for Compiling Dynamic Megakernel
di: Jin, Hongyi, et al.
Pubblicazione: (2026)
di: Jin, Hongyi, et al.
Pubblicazione: (2026)
Documenti analoghi
-
DeepCompile: A Compiler-Driven Approach to Optimizing Distributed Deep Learning Training
di: Tanaka, Masahiro, et al.
Pubblicazione: (2025) -
Flint: Compiler Enabled Cluster-Free Design Space Exploration for Distributed ML
di: Yoo, Jinsun, et al.
Pubblicazione: (2026) -
Efficient Parallel Compilation and Profiling of Quantum Circuits at Large Scales
di: Moore, Jane, et al.
Pubblicazione: (2026) -
LAPIS: A Performance Portable, High Productivity Compiler Framework
di: Kelley, Brian, et al.
Pubblicazione: (2025) -
Zen-Attention: A Compiler Framework for Dynamic Attention Folding on AMD NPUs
di: Deshmukh, Aadesh, et al.
Pubblicazione: (2025)