Stable Code Technical Report

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Pinnaparaju, Nikhil, Adithyan, Reshinth, Phung, Duy, Tow, Jonathan, Baicoianu, James, Datta, Ashish, Zhuravinskyi, Maksym, Mahan, Dakota, Bellagente, Marco, Riquelme, Carlos, Cooper, Nathan
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929299603324928
author Pinnaparaju, Nikhil
Adithyan, Reshinth
Phung, Duy
Tow, Jonathan
Baicoianu, James
Datta, Ashish
Zhuravinskyi, Maksym
Mahan, Dakota
Bellagente, Marco
Riquelme, Carlos
Cooper, Nathan
author_facet Pinnaparaju, Nikhil
Adithyan, Reshinth
Phung, Duy
Tow, Jonathan
Baicoianu, James
Datta, Ashish
Zhuravinskyi, Maksym
Mahan, Dakota
Bellagente, Marco
Riquelme, Carlos
Cooper, Nathan
contents We introduce Stable Code, the first in our new-generation of code language models series, which serves as a general-purpose base code language model targeting code completion, reasoning, math, and other software engineering-based tasks. Additionally, we introduce an instruction variant named Stable Code Instruct that allows conversing with the model in a natural chat interface for performing question-answering and instruction-based tasks. In this technical report, we detail the data and training procedure leading to both models. Their weights are available via Hugging Face for anyone to download and use at https://huggingface.co/stabilityai/stable-code-3b and https://huggingface.co/stabilityai/stable-code-instruct-3b. This report contains thorough evaluations of the models, including multilingual programming benchmarks, and the MT benchmark focusing on multi-turn dialogues. At the time of its release, Stable Code is the state-of-the-art open model under 3B parameters and even performs comparably to larger models of sizes 7 billion and 15 billion parameters on the popular Multi-PL benchmark. Stable Code Instruct also exhibits state-of-the-art performance on the MT-Bench coding tasks and on Multi-PL completion compared to other instruction tuned models. Given its appealing small size, we also provide throughput measurements on a number of edge devices. In addition, we open source several quantized checkpoints and provide their performance metrics compared to the original model.
format Preprint
id arxiv_https___arxiv_org_abs_2404_01226
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Stable Code Technical Report
Pinnaparaju, Nikhil
Adithyan, Reshinth
Phung, Duy
Tow, Jonathan
Baicoianu, James
Datta, Ashish
Zhuravinskyi, Maksym
Mahan, Dakota
Bellagente, Marco
Riquelme, Carlos
Cooper, Nathan
Computation and Language
We introduce Stable Code, the first in our new-generation of code language models series, which serves as a general-purpose base code language model targeting code completion, reasoning, math, and other software engineering-based tasks. Additionally, we introduce an instruction variant named Stable Code Instruct that allows conversing with the model in a natural chat interface for performing question-answering and instruction-based tasks. In this technical report, we detail the data and training procedure leading to both models. Their weights are available via Hugging Face for anyone to download and use at https://huggingface.co/stabilityai/stable-code-3b and https://huggingface.co/stabilityai/stable-code-instruct-3b. This report contains thorough evaluations of the models, including multilingual programming benchmarks, and the MT benchmark focusing on multi-turn dialogues. At the time of its release, Stable Code is the state-of-the-art open model under 3B parameters and even performs comparably to larger models of sizes 7 billion and 15 billion parameters on the popular Multi-PL benchmark. Stable Code Instruct also exhibits state-of-the-art performance on the MT-Bench coding tasks and on Multi-PL completion compared to other instruction tuned models. Given its appealing small size, we also provide throughput measurements on a number of edge devices. In addition, we open source several quantized checkpoints and provide their performance metrics compared to the original model.
title Stable Code Technical Report
topic Computation and Language
url https://arxiv.org/abs/2404.01226