SpikingBrain: Spiking Brain-inspired Large Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Pan, Yuqi, Feng, Yupeng, Zhuang, Jinghao, Ding, Siyu, Xu, Han, Liu, Zehao, Sun, Bohan, Chou, Yuhong, Qiu, Xuerui, Deng, Anlin, Hu, Anjie, Wang, Shurong, Zhou, Peng, Yao, Man, Wu, Jibin, Yang, Jian, Sun, Guoliang, Xu, Bo, Li, Guoqi
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910199742201856
author Pan, Yuqi
Feng, Yupeng
Zhuang, Jinghao
Ding, Siyu
Xu, Han
Liu, Zehao
Sun, Bohan
Chou, Yuhong
Qiu, Xuerui
Deng, Anlin
Hu, Anjie
Wang, Shurong
Zhou, Peng
Yao, Man
Wu, Jibin
Yang, Jian
Sun, Guoliang
Xu, Bo
Li, Guoqi
author_facet Pan, Yuqi
Feng, Yupeng
Zhuang, Jinghao
Ding, Siyu
Xu, Han
Liu, Zehao
Sun, Bohan
Chou, Yuhong
Qiu, Xuerui
Deng, Anlin
Hu, Anjie
Wang, Shurong
Zhou, Peng
Yao, Man
Wu, Jibin
Yang, Jian
Sun, Guoliang
Xu, Bo
Li, Guoqi
contents Mainstream Transformer-based large language models face major efficiency bottlenecks: training computation scales quadratically with sequence length, and inference memory grows linearly, limiting long-context processing. Building large models on non-NVIDIA platforms also poses challenges for stable and efficient training. To address this, we introduce SpikingBrain, a family of brain-inspired models designed for efficient long-context training and inference. SpikingBrain leverages the MetaX GPU cluster and focuses on three aspects: (1) Model Architecture: linear and hybrid-linear attention architectures with adaptive spiking neurons; (2) Algorithmic Optimizations: an efficient, conversion-based training pipeline and a dedicated spike coding framework; (3) System Engineering: customized training frameworks, operator libraries, and parallelism strategies tailored to MetaX hardware. Using these techniques, we develop two models: SpikingBrain-7B, a linear LLM, and SpikingBrain-76B, a hybrid-linear MoE LLM. These models demonstrate the feasibility of large-scale LLM development on non-NVIDIA platforms, and training remains stable for weeks on hundreds of MetaX GPUs with Model FLOPs Utilization at expected levels. SpikingBrain achieves performance comparable to open-source Transformer baselines while using only about 150B tokens for continual pre-training. Our models also significantly improve long-context efficiency and deliver inference with (partially) constant memory and event-driven spiking behavior. For example, SpikingBrain-7B attains over 100x speedup in Time to First Token for 4M-token sequences. Furthermore, the proposed spiking scheme achieves 69.15 percent sparsity, enabling low-power operation. Overall, this work demonstrates the potential of brain-inspired mechanisms to drive the next generation of efficient and scalable large model design.
format Preprint
id arxiv_https___arxiv_org_abs_2509_05276
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SpikingBrain: Spiking Brain-inspired Large Models
Pan, Yuqi
Feng, Yupeng
Zhuang, Jinghao
Ding, Siyu
Xu, Han
Liu, Zehao
Sun, Bohan
Chou, Yuhong
Qiu, Xuerui
Deng, Anlin
Hu, Anjie
Wang, Shurong
Zhou, Peng
Yao, Man
Wu, Jibin
Yang, Jian
Sun, Guoliang
Xu, Bo
Li, Guoqi
Machine Learning
Artificial Intelligence
Computation and Language
Mainstream Transformer-based large language models face major efficiency bottlenecks: training computation scales quadratically with sequence length, and inference memory grows linearly, limiting long-context processing. Building large models on non-NVIDIA platforms also poses challenges for stable and efficient training. To address this, we introduce SpikingBrain, a family of brain-inspired models designed for efficient long-context training and inference. SpikingBrain leverages the MetaX GPU cluster and focuses on three aspects: (1) Model Architecture: linear and hybrid-linear attention architectures with adaptive spiking neurons; (2) Algorithmic Optimizations: an efficient, conversion-based training pipeline and a dedicated spike coding framework; (3) System Engineering: customized training frameworks, operator libraries, and parallelism strategies tailored to MetaX hardware. Using these techniques, we develop two models: SpikingBrain-7B, a linear LLM, and SpikingBrain-76B, a hybrid-linear MoE LLM. These models demonstrate the feasibility of large-scale LLM development on non-NVIDIA platforms, and training remains stable for weeks on hundreds of MetaX GPUs with Model FLOPs Utilization at expected levels. SpikingBrain achieves performance comparable to open-source Transformer baselines while using only about 150B tokens for continual pre-training. Our models also significantly improve long-context efficiency and deliver inference with (partially) constant memory and event-driven spiking behavior. For example, SpikingBrain-7B attains over 100x speedup in Time to First Token for 4M-token sequences. Furthermore, the proposed spiking scheme achieves 69.15 percent sparsity, enabling low-power operation. Overall, this work demonstrates the potential of brain-inspired mechanisms to drive the next generation of efficient and scalable large model design.
title SpikingBrain: Spiking Brain-inspired Large Models
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2509.05276