ChuXin: 1.6B Technical Report
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916239119482880 |
|---|---|
| author | Zhuang, Xiaomin Jiang, Yufan He, Qiaozhi Wu, Zhihua |
| author_facet | Zhuang, Xiaomin Jiang, Yufan He, Qiaozhi Wu, Zhihua |
| contents | In this report, we present ChuXin, an entirely open-source language model with a size of 1.6 billion parameters. Unlike the majority of works that only open-sourced the model weights and architecture, we have made everything needed to train a model available, including the training data, the training process, and the evaluation code. Our goal is to empower and strengthen the open research community, fostering transparency and enabling a new wave of innovation in the field of language modeling. Furthermore, we extend the context length to 1M tokens through lightweight continual pretraining and demonstrate strong needle-in-a-haystack retrieval performance. The weights for both models are available at Hugging Face to download and use. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2405_04828 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | ChuXin: 1.6B Technical Report Zhuang, Xiaomin Jiang, Yufan He, Qiaozhi Wu, Zhihua Computation and Language In this report, we present ChuXin, an entirely open-source language model with a size of 1.6 billion parameters. Unlike the majority of works that only open-sourced the model weights and architecture, we have made everything needed to train a model available, including the training data, the training process, and the evaluation code. Our goal is to empower and strengthen the open research community, fostering transparency and enabling a new wave of innovation in the field of language modeling. Furthermore, we extend the context length to 1M tokens through lightweight continual pretraining and demonstrate strong needle-in-a-haystack retrieval performance. The weights for both models are available at Hugging Face to download and use. |
| title | ChuXin: 1.6B Technical Report |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2405.04828 |