ARCHE3-7B: A Hierarchical Mixture-of-Experts Architecture with Foundation Curriculum Training

Fuente: Zenodo
Gespeichert in:
Bibliographische Detailangaben
1. Verfasser: Osovskoi, Ilya
Format: Recurso digital
Sprache:Englisch
Veröffentlicht: Zenodo 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866902103658594304
author Osovskoi, Ilya
author_facet Osovskoi, Ilya
contents <p>ARCHE3-7B is not just another language model. It’s a fundamental shift from "guessing the next word" to "understanding how the world works." This manifesto outlines the architecture and philosophy behind a system designed to be the "brain" for the next generation of robotics.</p> <p>Key Breakthroughs:</p> <p>FCT (Foundation Curriculum Training): We don't start with text. We start with the "Source Code of Reality." The model first internalizes 290 structural patterns (logic, physics, systems, causality) to build a rational scaffold before it ever reads a single sentence.</p> <p>Split Dense Core: A dual-stage brain architecture. An Input Core for perception and a Fusion Core for decision-making. It mimics the human brain’s separation of sensory processing and cognitive synthesis.</p> <p>HMoE (20,480 Experts): A massive swarm of specialized experts stored on-disk. This allows a 7B-parameter model to run on just a few gigabytes of RAM, making high-level intelligence possible on consumer-grade hardware and edge devices.</p> <p>Rational Ethics: No hard-coded word filters or "censor" layers. The model uses an internal Objective System (Survival, Human Protection, Efficiency) to evaluate its own actions through logic, not just rules.</p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_18738608
institution Zenodo
language eng
publishDate 2026
publisher Zenodo
record_format zenodo
spellingShingle ARCHE3-7B: A Hierarchical Mixture-of-Experts Architecture with Foundation Curriculum Training
Osovskoi, Ilya
World Model
Sparse Attention
ARCHE3-7B
MoE
<p>ARCHE3-7B is not just another language model. It’s a fundamental shift from "guessing the next word" to "understanding how the world works." This manifesto outlines the architecture and philosophy behind a system designed to be the "brain" for the next generation of robotics.</p> <p>Key Breakthroughs:</p> <p>FCT (Foundation Curriculum Training): We don't start with text. We start with the "Source Code of Reality." The model first internalizes 290 structural patterns (logic, physics, systems, causality) to build a rational scaffold before it ever reads a single sentence.</p> <p>Split Dense Core: A dual-stage brain architecture. An Input Core for perception and a Fusion Core for decision-making. It mimics the human brain’s separation of sensory processing and cognitive synthesis.</p> <p>HMoE (20,480 Experts): A massive swarm of specialized experts stored on-disk. This allows a 7B-parameter model to run on just a few gigabytes of RAM, making high-level intelligence possible on consumer-grade hardware and edge devices.</p> <p>Rational Ethics: No hard-coded word filters or "censor" layers. The model uses an internal Objective System (Survival, Human Protection, Efficiency) to evaluate its own actions through logic, not just rules.</p>
title ARCHE3-7B: A Hierarchical Mixture-of-Experts Architecture with Foundation Curriculum Training
topic World Model
Sparse Attention
ARCHE3-7B
MoE
url https://doi.org/10.5281/zenodo.18738608