Aquila2 Technical Report

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Bo-Wen, Wang, Liangdong, Li, Jijie, Gu, Shuhao, Wu, Xinya, Zhang, Zhengduo, Gao, Boyan, Ao, Yulong, Liu, Guang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909286975668224
author Zhang, Bo-Wen
Wang, Liangdong
Li, Jijie
Gu, Shuhao
Wu, Xinya
Zhang, Zhengduo
Gao, Boyan
Ao, Yulong
Liu, Guang
author_facet Zhang, Bo-Wen
Wang, Liangdong
Li, Jijie
Gu, Shuhao
Wu, Xinya
Zhang, Zhengduo
Gao, Boyan
Ao, Yulong
Liu, Guang
contents This paper introduces the Aquila2 series, which comprises a wide range of bilingual models with parameter sizes of 7, 34, and 70 billion. These models are trained based on an innovative framework named HeuriMentor (HM), which offers real-time insights into model convergence and enhances the training process and data management. The HM System, comprising the Adaptive Training Engine (ATE), Training State Monitor (TSM), and Data Management Unit (DMU), allows for precise monitoring of the model's training progress and enables efficient optimization of data distribution, thereby enhancing training effectiveness. Extensive evaluations show that the Aquila2 model series performs comparably well on both English and Chinese benchmarks. Specifically, Aquila2-34B demonstrates only a slight decrease in performance when quantized to Int4. Furthermore, we have made our training code (https://github.com/FlagOpen/FlagScale) and model weights (https://github.com/FlagAI-Open/Aquila2) publicly available to support ongoing research and the development of applications.
format Preprint
id arxiv_https___arxiv_org_abs_2408_07410
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Aquila2 Technical Report
Zhang, Bo-Wen
Wang, Liangdong
Li, Jijie
Gu, Shuhao
Wu, Xinya
Zhang, Zhengduo
Gao, Boyan
Ao, Yulong
Liu, Guang
Computation and Language
This paper introduces the Aquila2 series, which comprises a wide range of bilingual models with parameter sizes of 7, 34, and 70 billion. These models are trained based on an innovative framework named HeuriMentor (HM), which offers real-time insights into model convergence and enhances the training process and data management. The HM System, comprising the Adaptive Training Engine (ATE), Training State Monitor (TSM), and Data Management Unit (DMU), allows for precise monitoring of the model's training progress and enables efficient optimization of data distribution, thereby enhancing training effectiveness. Extensive evaluations show that the Aquila2 model series performs comparably well on both English and Chinese benchmarks. Specifically, Aquila2-34B demonstrates only a slight decrease in performance when quantized to Int4. Furthermore, we have made our training code (https://github.com/FlagOpen/FlagScale) and model weights (https://github.com/FlagAI-Open/Aquila2) publicly available to support ongoing research and the development of applications.
title Aquila2 Technical Report
topic Computation and Language
url https://arxiv.org/abs/2408.07410