MonoCoder: Domain-Specific Code Language Model for HPC Codes and Tasks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kadosh, Tal, Hasabnis, Niranjan, Vo, Vy A., Schneider, Nadav, Krien, Neva, Capota, Mihai, Wasay, Abdul, Ahmed, Nesreen, Willke, Ted, Tamir, Guy, Pinter, Yuval, Mattson, Timothy, Oren, Gal
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916401820729344
author Kadosh, Tal
Hasabnis, Niranjan
Vo, Vy A.
Schneider, Nadav
Krien, Neva
Capota, Mihai
Wasay, Abdul
Ahmed, Nesreen
Willke, Ted
Tamir, Guy
Pinter, Yuval
Mattson, Timothy
Oren, Gal
author_facet Kadosh, Tal
Hasabnis, Niranjan
Vo, Vy A.
Schneider, Nadav
Krien, Neva
Capota, Mihai
Wasay, Abdul
Ahmed, Nesreen
Willke, Ted
Tamir, Guy
Pinter, Yuval
Mattson, Timothy
Oren, Gal
contents With easier access to powerful compute resources, there is a growing trend in AI for software development to develop large language models (LLMs) to address a variety of programming tasks. Even LLMs applied to tasks from the high-performance computing (HPC) domain are huge in size and demand expensive compute resources for training. This is partly because LLMs for HPC tasks are obtained by finetuning existing LLMs that support several natural and/or programming languages. We found this design choice confusing - why do we need LLMs trained on natural languages and programming languages unrelated to HPC for HPC-specific tasks? In this line of work, we aim to question choices made by existing LLMs by developing smaller language models (LMs) for specific domains - we call them domain-specific LMs. Specifically, we start with HPC as a domain and build an HPC-specific LM, named MonoCoder, which is orders of magnitude smaller than existing LMs but delivers better performance on non-HPC and HPC codes. Specifically, we pre-trained MonoCoder on an HPC-specific dataset (named HPCorpus) of C and C++ programs mined from GitHub. We evaluated the performance of MonoCoder against state-of-the-art multi-lingual LLMs. Results demonstrate that MonoCoder, although much smaller than existing LMs, outperforms other LLMs on normalized-perplexity tests (in relation to model size) while also delivering competing CodeBLEU scores for high-performance and parallel code generations. In other words, results suggest that MonoCoder understands HPC code better than state-of-the-art LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2312_13322
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle MonoCoder: Domain-Specific Code Language Model for HPC Codes and Tasks
Kadosh, Tal
Hasabnis, Niranjan
Vo, Vy A.
Schneider, Nadav
Krien, Neva
Capota, Mihai
Wasay, Abdul
Ahmed, Nesreen
Willke, Ted
Tamir, Guy
Pinter, Yuval
Mattson, Timothy
Oren, Gal
Programming Languages
Artificial Intelligence
Machine Learning
Software Engineering
With easier access to powerful compute resources, there is a growing trend in AI for software development to develop large language models (LLMs) to address a variety of programming tasks. Even LLMs applied to tasks from the high-performance computing (HPC) domain are huge in size and demand expensive compute resources for training. This is partly because LLMs for HPC tasks are obtained by finetuning existing LLMs that support several natural and/or programming languages. We found this design choice confusing - why do we need LLMs trained on natural languages and programming languages unrelated to HPC for HPC-specific tasks? In this line of work, we aim to question choices made by existing LLMs by developing smaller language models (LMs) for specific domains - we call them domain-specific LMs. Specifically, we start with HPC as a domain and build an HPC-specific LM, named MonoCoder, which is orders of magnitude smaller than existing LMs but delivers better performance on non-HPC and HPC codes. Specifically, we pre-trained MonoCoder on an HPC-specific dataset (named HPCorpus) of C and C++ programs mined from GitHub. We evaluated the performance of MonoCoder against state-of-the-art multi-lingual LLMs. Results demonstrate that MonoCoder, although much smaller than existing LMs, outperforms other LLMs on normalized-perplexity tests (in relation to model size) while also delivering competing CodeBLEU scores for high-performance and parallel code generations. In other words, results suggest that MonoCoder understands HPC code better than state-of-the-art LLMs.
title MonoCoder: Domain-Specific Code Language Model for HPC Codes and Tasks
topic Programming Languages
Artificial Intelligence
Machine Learning
Software Engineering
url https://arxiv.org/abs/2312.13322