CCI3.0-HQ: a large-scale Chinese dataset of high quality designed for pre-training large language models
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866914988056117248 |
|---|---|
| author | Wang, Liangdong Zhang, Bo-Wen Wu, Chengwei Zhao, Hanyu Shi, Xiaofeng Gu, Shuhao Li, Jijie Ma, Quanyue Pan, TengFei Liu, Guang |
| author_facet | Wang, Liangdong Zhang, Bo-Wen Wu, Chengwei Zhao, Hanyu Shi, Xiaofeng Gu, Shuhao Li, Jijie Ma, Quanyue Pan, TengFei Liu, Guang |
| contents | We present CCI3.0-HQ (https://huggingface.co/datasets/BAAI/CCI3-HQ), a high-quality 500GB subset of the Chinese Corpora Internet 3.0 (CCI3.0)(https://huggingface.co/datasets/BAAI/CCI3-Data), developed using a novel two-stage hybrid filtering pipeline that significantly enhances data quality. To evaluate its effectiveness, we trained a 0.5B parameter model from scratch on 100B tokens across various datasets, achieving superior performance on 10 benchmarks in a zero-shot setting compared to CCI3.0, SkyPile, and WanjuanV1. The high-quality filtering process effectively distills the capabilities of the Qwen2-72B-instruct model into a compact 0.5B model, attaining optimal F1 scores for Chinese web data classification. We believe this open-access dataset will facilitate broader access to high-quality language models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_18505 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | CCI3.0-HQ: a large-scale Chinese dataset of high quality designed for pre-training large language models Wang, Liangdong Zhang, Bo-Wen Wu, Chengwei Zhao, Hanyu Shi, Xiaofeng Gu, Shuhao Li, Jijie Ma, Quanyue Pan, TengFei Liu, Guang Computation and Language We present CCI3.0-HQ (https://huggingface.co/datasets/BAAI/CCI3-HQ), a high-quality 500GB subset of the Chinese Corpora Internet 3.0 (CCI3.0)(https://huggingface.co/datasets/BAAI/CCI3-Data), developed using a novel two-stage hybrid filtering pipeline that significantly enhances data quality. To evaluate its effectiveness, we trained a 0.5B parameter model from scratch on 100B tokens across various datasets, achieving superior performance on 10 benchmarks in a zero-shot setting compared to CCI3.0, SkyPile, and WanjuanV1. The high-quality filtering process effectively distills the capabilities of the Qwen2-72B-instruct model into a compact 0.5B model, attaining optimal F1 scores for Chinese web data classification. We believe this open-access dataset will facilitate broader access to high-quality language models. |
| title | CCI3.0-HQ: a large-scale Chinese dataset of high quality designed for pre-training large language models |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2410.18505 |