Arcee Trinity Large Technical Report

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Singh, Varun, Krauss, Lucas, Jaghouar, Sami, Sirovatka, Matej, Goddard, Charles, Obied, Fares, Ong, Jack Min, Straube, Jannik, Fern, Harley, Aria, Stewart, Conner, Kealty, Colin, Panahi, Maziyar, Kirsten, Simon, Deshpande, Anushka, Vij, Anneketh, Bresnu, Arthur, Veldurthi, Pranav, Ravishankar, Raghav, Bishnoi, Hardik, Team, DatologyAI, Team, Arcee AI, Team, Prime Intellect, McQuade, Mark, Hagemann, Johannes, Atkins, Lucas
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915806360633344
author Singh, Varun
Krauss, Lucas
Jaghouar, Sami
Sirovatka, Matej
Goddard, Charles
Obied, Fares
Ong, Jack Min
Straube, Jannik
Fern
Harley, Aria
Stewart, Conner
Kealty, Colin
Panahi, Maziyar
Kirsten, Simon
Deshpande, Anushka
Vij, Anneketh
Bresnu, Arthur
Veldurthi, Pranav
Ravishankar, Raghav
Bishnoi, Hardik
Team, DatologyAI
Team, Arcee AI
Team, Prime Intellect
McQuade, Mark
Hagemann, Johannes
Atkins, Lucas
author_facet Singh, Varun
Krauss, Lucas
Jaghouar, Sami
Sirovatka, Matej
Goddard, Charles
Obied, Fares
Ong, Jack Min
Straube, Jannik
Fern
Harley, Aria
Stewart, Conner
Kealty, Colin
Panahi, Maziyar
Kirsten, Simon
Deshpande, Anushka
Vij, Anneketh
Bresnu, Arthur
Veldurthi, Pranav
Ravishankar, Raghav
Bishnoi, Hardik
Team, DatologyAI
Team, Arcee AI
Team, Prime Intellect
McQuade, Mark
Hagemann, Johannes
Atkins, Lucas
contents We present the technical report for Arcee Trinity Large, a sparse Mixture-of-Experts model with 400B total parameters and 13B activated per token. Additionally, we report on Trinity Nano and Trinity Mini, with Trinity Nano having 6B total parameters with 1B activated per token, Trinity Mini having 26B total parameters with 3B activated per token. The models' modern architecture includes interleaved local and global attention, gated attention, depth-scaled sandwich norm, and sigmoid routing for Mixture-of-Experts. For Trinity Large, we also introduce a new MoE load balancing strategy titled Soft-clamped Momentum Expert Bias Updates (SMEBU). We train the models using the Muon optimizer. All three models completed training with zero loss spikes. Trinity Nano and Trinity Mini were pre-trained on 10 trillion tokens, and Trinity Large was pre-trained on 17 trillion tokens. The model checkpoints are available at https://huggingface.co/arcee-ai.
format Preprint
id arxiv_https___arxiv_org_abs_2602_17004
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Arcee Trinity Large Technical Report
Singh, Varun
Krauss, Lucas
Jaghouar, Sami
Sirovatka, Matej
Goddard, Charles
Obied, Fares
Ong, Jack Min
Straube, Jannik
Fern
Harley, Aria
Stewart, Conner
Kealty, Colin
Panahi, Maziyar
Kirsten, Simon
Deshpande, Anushka
Vij, Anneketh
Bresnu, Arthur
Veldurthi, Pranav
Ravishankar, Raghav
Bishnoi, Hardik
Team, DatologyAI
Team, Arcee AI
Team, Prime Intellect
McQuade, Mark
Hagemann, Johannes
Atkins, Lucas
Machine Learning
Computation and Language
We present the technical report for Arcee Trinity Large, a sparse Mixture-of-Experts model with 400B total parameters and 13B activated per token. Additionally, we report on Trinity Nano and Trinity Mini, with Trinity Nano having 6B total parameters with 1B activated per token, Trinity Mini having 26B total parameters with 3B activated per token. The models' modern architecture includes interleaved local and global attention, gated attention, depth-scaled sandwich norm, and sigmoid routing for Mixture-of-Experts. For Trinity Large, we also introduce a new MoE load balancing strategy titled Soft-clamped Momentum Expert Bias Updates (SMEBU). We train the models using the Muon optimizer. All three models completed training with zero loss spikes. Trinity Nano and Trinity Mini were pre-trained on 10 trillion tokens, and Trinity Large was pre-trained on 17 trillion tokens. The model checkpoints are available at https://huggingface.co/arcee-ai.
title Arcee Trinity Large Technical Report
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2602.17004