INTELLECT-2: A Reasoning Model Trained Through Globally Decentralized Reinforcement Learning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Prime Intellect Team, Jaghouar, Sami, Mattern, Justus, Ong, Jack Min, Straube, Jannik, Basra, Manveer, Pazdera, Aaron, Thaman, Kushal, Di Ferrante, Matthew, Gabriel, Felix, Obeid, Fares, Erdem, Kemal, Keiblinger, Michael, Hagemann, Johannes
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918017628110848
author Prime Intellect Team
Jaghouar, Sami
Mattern, Justus
Ong, Jack Min
Straube, Jannik
Basra, Manveer
Pazdera, Aaron
Thaman, Kushal
Di Ferrante, Matthew
Gabriel, Felix
Obeid, Fares
Erdem, Kemal
Keiblinger, Michael
Hagemann, Johannes
author_facet Prime Intellect Team
Jaghouar, Sami
Mattern, Justus
Ong, Jack Min
Straube, Jannik
Basra, Manveer
Pazdera, Aaron
Thaman, Kushal
Di Ferrante, Matthew
Gabriel, Felix
Obeid, Fares
Erdem, Kemal
Keiblinger, Michael
Hagemann, Johannes
contents We introduce INTELLECT-2, the first globally distributed reinforcement learning (RL) training run of a 32 billion parameter language model. Unlike traditional centralized training efforts, INTELLECT-2 trains a reasoning model using fully asynchronous RL across a dynamic, heterogeneous swarm of permissionless compute contributors. To enable a training run with this unique infrastructure, we built various components from scratch: we introduce PRIME-RL, our training framework purpose-built for distributed asynchronous reinforcement learning, based on top of novel components such as TOPLOC, which verifies rollouts from untrusted inference workers, and SHARDCAST, which efficiently broadcasts policy weights from training nodes to inference workers. Beyond infrastructure components, we propose modifications to the standard GRPO training recipe and data filtering techniques that were crucial to achieve training stability and ensure that our model successfully learned its training objective, thus improving upon QwQ-32B, the state of the art reasoning model in the 32B parameter range. We open-source INTELLECT-2 along with all of our code and data, hoping to encourage and enable more open research in the field of decentralized training.
format Preprint
id arxiv_https___arxiv_org_abs_2505_07291
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle INTELLECT-2: A Reasoning Model Trained Through Globally Decentralized Reinforcement Learning
Prime Intellect Team
Jaghouar, Sami
Mattern, Justus
Ong, Jack Min
Straube, Jannik
Basra, Manveer
Pazdera, Aaron
Thaman, Kushal
Di Ferrante, Matthew
Gabriel, Felix
Obeid, Fares
Erdem, Kemal
Keiblinger, Michael
Hagemann, Johannes
Machine Learning
Distributed, Parallel, and Cluster Computing
We introduce INTELLECT-2, the first globally distributed reinforcement learning (RL) training run of a 32 billion parameter language model. Unlike traditional centralized training efforts, INTELLECT-2 trains a reasoning model using fully asynchronous RL across a dynamic, heterogeneous swarm of permissionless compute contributors. To enable a training run with this unique infrastructure, we built various components from scratch: we introduce PRIME-RL, our training framework purpose-built for distributed asynchronous reinforcement learning, based on top of novel components such as TOPLOC, which verifies rollouts from untrusted inference workers, and SHARDCAST, which efficiently broadcasts policy weights from training nodes to inference workers. Beyond infrastructure components, we propose modifications to the standard GRPO training recipe and data filtering techniques that were crucial to achieve training stability and ensure that our model successfully learned its training objective, thus improving upon QwQ-32B, the state of the art reasoning model in the 32B parameter range. We open-source INTELLECT-2 along with all of our code and data, hoping to encourage and enable more open research in the field of decentralized training.
title INTELLECT-2: A Reasoning Model Trained Through Globally Decentralized Reinforcement Learning
topic Machine Learning
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2505.07291