MojoBench: Language Modeling and Benchmarks for Mojo

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Raihan, Nishat, Santos, Joanna C. S., Zampieri, Marcos
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916449527791616
author Raihan, Nishat
Santos, Joanna C. S.
Zampieri, Marcos
author_facet Raihan, Nishat
Santos, Joanna C. S.
Zampieri, Marcos
contents The recently introduced Mojo programming language (PL) by Modular, has received significant attention in the scientific community due to its claimed significant speed boost over Python. Despite advancements in code Large Language Models (LLMs) across various PLs, Mojo remains unexplored in this context. To address this gap, we introduce MojoBench, the first framework for Mojo code generation. MojoBench includes HumanEval-Mojo, a benchmark dataset designed for evaluating code LLMs on Mojo, and Mojo-Coder, the first LLM pretrained and finetuned for Mojo code generation, which supports instructions in 5 natural languages (NLs). Our results show that Mojo-Coder achieves a 30-35% performance improvement over leading models like GPT-4o and Claude-3.5-Sonnet. Furthermore, we provide insights into LLM behavior with underrepresented and unseen PLs, offering potential strategies for enhancing model adaptability. MojoBench contributes to our understanding of LLM capabilities and limitations in emerging programming paradigms fostering more robust code generation systems.
format Preprint
id arxiv_https___arxiv_org_abs_2410_17736
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MojoBench: Language Modeling and Benchmarks for Mojo
Raihan, Nishat
Santos, Joanna C. S.
Zampieri, Marcos
Computation and Language
The recently introduced Mojo programming language (PL) by Modular, has received significant attention in the scientific community due to its claimed significant speed boost over Python. Despite advancements in code Large Language Models (LLMs) across various PLs, Mojo remains unexplored in this context. To address this gap, we introduce MojoBench, the first framework for Mojo code generation. MojoBench includes HumanEval-Mojo, a benchmark dataset designed for evaluating code LLMs on Mojo, and Mojo-Coder, the first LLM pretrained and finetuned for Mojo code generation, which supports instructions in 5 natural languages (NLs). Our results show that Mojo-Coder achieves a 30-35% performance improvement over leading models like GPT-4o and Claude-3.5-Sonnet. Furthermore, we provide insights into LLM behavior with underrepresented and unseen PLs, offering potential strategies for enhancing model adaptability. MojoBench contributes to our understanding of LLM capabilities and limitations in emerging programming paradigms fostering more robust code generation systems.
title MojoBench: Language Modeling and Benchmarks for Mojo
topic Computation and Language
url https://arxiv.org/abs/2410.17736