Scaling Granite Code Models to 128K Context

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Stallone, Matt, Saxena, Vaibhav, Karlinsky, Leonid, McGinn, Bridget, Bula, Tim, Mishra, Mayank, Soria, Adriana Meza, Zhang, Gaoyuan, Prasad, Aditya, Shen, Yikang, Surendran, Saptha, Guttula, Shanmukha, Patel, Hima, Selvam, Parameswaran, Dang, Xuan-Hong, Koyfman, Yan, Sood, Atin, Feris, Rogerio, Desai, Nirmit, Cox, David D., Puri, Ruchir, Panda, Rameswar
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910534042910720
author Stallone, Matt
Saxena, Vaibhav
Karlinsky, Leonid
McGinn, Bridget
Bula, Tim
Mishra, Mayank
Soria, Adriana Meza
Zhang, Gaoyuan
Prasad, Aditya
Shen, Yikang
Surendran, Saptha
Guttula, Shanmukha
Patel, Hima
Selvam, Parameswaran
Dang, Xuan-Hong
Koyfman, Yan
Sood, Atin
Feris, Rogerio
Desai, Nirmit
Cox, David D.
Puri, Ruchir
Panda, Rameswar
author_facet Stallone, Matt
Saxena, Vaibhav
Karlinsky, Leonid
McGinn, Bridget
Bula, Tim
Mishra, Mayank
Soria, Adriana Meza
Zhang, Gaoyuan
Prasad, Aditya
Shen, Yikang
Surendran, Saptha
Guttula, Shanmukha
Patel, Hima
Selvam, Parameswaran
Dang, Xuan-Hong
Koyfman, Yan
Sood, Atin
Feris, Rogerio
Desai, Nirmit
Cox, David D.
Puri, Ruchir
Panda, Rameswar
contents This paper introduces long-context Granite code models that support effective context windows of up to 128K tokens. Our solution for scaling context length of Granite 3B/8B code models from 2K/4K to 128K consists of a light-weight continual pretraining by gradually increasing its RoPE base frequency with repository-level file packing and length-upsampled long-context data. Additionally, we also release instruction-tuned models with long-context support which are derived by further finetuning the long context base models on a mix of permissively licensed short and long-context instruction-response pairs. While comparing to the original short-context Granite code models, our long-context models achieve significant improvements on long-context tasks without any noticeable performance degradation on regular code completion benchmarks (e.g., HumanEval). We release all our long-context Granite code models under an Apache 2.0 license for both research and commercial use.
format Preprint
id arxiv_https___arxiv_org_abs_2407_13739
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Scaling Granite Code Models to 128K Context
Stallone, Matt
Saxena, Vaibhav
Karlinsky, Leonid
McGinn, Bridget
Bula, Tim
Mishra, Mayank
Soria, Adriana Meza
Zhang, Gaoyuan
Prasad, Aditya
Shen, Yikang
Surendran, Saptha
Guttula, Shanmukha
Patel, Hima
Selvam, Parameswaran
Dang, Xuan-Hong
Koyfman, Yan
Sood, Atin
Feris, Rogerio
Desai, Nirmit
Cox, David D.
Puri, Ruchir
Panda, Rameswar
Artificial Intelligence
Computation and Language
Software Engineering
This paper introduces long-context Granite code models that support effective context windows of up to 128K tokens. Our solution for scaling context length of Granite 3B/8B code models from 2K/4K to 128K consists of a light-weight continual pretraining by gradually increasing its RoPE base frequency with repository-level file packing and length-upsampled long-context data. Additionally, we also release instruction-tuned models with long-context support which are derived by further finetuning the long context base models on a mix of permissively licensed short and long-context instruction-response pairs. While comparing to the original short-context Granite code models, our long-context models achieve significant improvements on long-context tasks without any noticeable performance degradation on regular code completion benchmarks (e.g., HumanEval). We release all our long-context Granite code models under an Apache 2.0 license for both research and commercial use.
title Scaling Granite Code Models to 128K Context
topic Artificial Intelligence
Computation and Language
Software Engineering
url https://arxiv.org/abs/2407.13739