Arcee Trinity Large Technical Report
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866915806360633344 |
|---|---|
| author | Singh, Varun Krauss, Lucas Jaghouar, Sami Sirovatka, Matej Goddard, Charles Obied, Fares Ong, Jack Min Straube, Jannik Fern Harley, Aria Stewart, Conner Kealty, Colin Panahi, Maziyar Kirsten, Simon Deshpande, Anushka Vij, Anneketh Bresnu, Arthur Veldurthi, Pranav Ravishankar, Raghav Bishnoi, Hardik Team, DatologyAI Team, Arcee AI Team, Prime Intellect McQuade, Mark Hagemann, Johannes Atkins, Lucas |
| author_facet | Singh, Varun Krauss, Lucas Jaghouar, Sami Sirovatka, Matej Goddard, Charles Obied, Fares Ong, Jack Min Straube, Jannik Fern Harley, Aria Stewart, Conner Kealty, Colin Panahi, Maziyar Kirsten, Simon Deshpande, Anushka Vij, Anneketh Bresnu, Arthur Veldurthi, Pranav Ravishankar, Raghav Bishnoi, Hardik Team, DatologyAI Team, Arcee AI Team, Prime Intellect McQuade, Mark Hagemann, Johannes Atkins, Lucas |
| contents | We present the technical report for Arcee Trinity Large, a sparse Mixture-of-Experts model with 400B total parameters and 13B activated per token. Additionally, we report on Trinity Nano and Trinity Mini, with Trinity Nano having 6B total parameters with 1B activated per token, Trinity Mini having 26B total parameters with 3B activated per token. The models' modern architecture includes interleaved local and global attention, gated attention, depth-scaled sandwich norm, and sigmoid routing for Mixture-of-Experts. For Trinity Large, we also introduce a new MoE load balancing strategy titled Soft-clamped Momentum Expert Bias Updates (SMEBU). We train the models using the Muon optimizer. All three models completed training with zero loss spikes. Trinity Nano and Trinity Mini were pre-trained on 10 trillion tokens, and Trinity Large was pre-trained on 17 trillion tokens. The model checkpoints are available at https://huggingface.co/arcee-ai. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_17004 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Arcee Trinity Large Technical Report Singh, Varun Krauss, Lucas Jaghouar, Sami Sirovatka, Matej Goddard, Charles Obied, Fares Ong, Jack Min Straube, Jannik Fern Harley, Aria Stewart, Conner Kealty, Colin Panahi, Maziyar Kirsten, Simon Deshpande, Anushka Vij, Anneketh Bresnu, Arthur Veldurthi, Pranav Ravishankar, Raghav Bishnoi, Hardik Team, DatologyAI Team, Arcee AI Team, Prime Intellect McQuade, Mark Hagemann, Johannes Atkins, Lucas Machine Learning Computation and Language We present the technical report for Arcee Trinity Large, a sparse Mixture-of-Experts model with 400B total parameters and 13B activated per token. Additionally, we report on Trinity Nano and Trinity Mini, with Trinity Nano having 6B total parameters with 1B activated per token, Trinity Mini having 26B total parameters with 3B activated per token. The models' modern architecture includes interleaved local and global attention, gated attention, depth-scaled sandwich norm, and sigmoid routing for Mixture-of-Experts. For Trinity Large, we also introduce a new MoE load balancing strategy titled Soft-clamped Momentum Expert Bias Updates (SMEBU). We train the models using the Muon optimizer. All three models completed training with zero loss spikes. Trinity Nano and Trinity Mini were pre-trained on 10 trillion tokens, and Trinity Large was pre-trained on 17 trillion tokens. The model checkpoints are available at https://huggingface.co/arcee-ai. |
| title | Arcee Trinity Large Technical Report |
| topic | Machine Learning Computation and Language |
| url | https://arxiv.org/abs/2602.17004 |