ARCHE3-7B: A Hierarchical Mixture-of-Experts Architecture with Foundation Curriculum Training
Fuente:
Zenodo
Gespeichert in:
| 1. Verfasser: | |
|---|---|
| Format: | Recurso digital |
| Sprache: | Englisch |
| Veröffentlicht: |
Zenodo
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866902103658594304 |
|---|---|
| author | Osovskoi, Ilya |
| author_facet | Osovskoi, Ilya |
| contents | <p>ARCHE3-7B is not just another language model. It’s a fundamental shift from "guessing the next word" to "understanding how the world works." This manifesto outlines the architecture and philosophy behind a system designed to be the "brain" for the next generation of robotics.</p> <p>Key Breakthroughs:</p> <p>FCT (Foundation Curriculum Training): We don't start with text. We start with the "Source Code of Reality." The model first internalizes 290 structural patterns (logic, physics, systems, causality) to build a rational scaffold before it ever reads a single sentence.</p> <p>Split Dense Core: A dual-stage brain architecture. An Input Core for perception and a Fusion Core for decision-making. It mimics the human brain’s separation of sensory processing and cognitive synthesis.</p> <p>HMoE (20,480 Experts): A massive swarm of specialized experts stored on-disk. This allows a 7B-parameter model to run on just a few gigabytes of RAM, making high-level intelligence possible on consumer-grade hardware and edge devices.</p> <p>Rational Ethics: No hard-coded word filters or "censor" layers. The model uses an internal Objective System (Survival, Human Protection, Efficiency) to evaluate its own actions through logic, not just rules.</p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_18738608 |
| institution | Zenodo |
| language | eng |
| publishDate | 2026 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | ARCHE3-7B: A Hierarchical Mixture-of-Experts Architecture with Foundation Curriculum Training Osovskoi, Ilya World Model Sparse Attention ARCHE3-7B MoE <p>ARCHE3-7B is not just another language model. It’s a fundamental shift from "guessing the next word" to "understanding how the world works." This manifesto outlines the architecture and philosophy behind a system designed to be the "brain" for the next generation of robotics.</p> <p>Key Breakthroughs:</p> <p>FCT (Foundation Curriculum Training): We don't start with text. We start with the "Source Code of Reality." The model first internalizes 290 structural patterns (logic, physics, systems, causality) to build a rational scaffold before it ever reads a single sentence.</p> <p>Split Dense Core: A dual-stage brain architecture. An Input Core for perception and a Fusion Core for decision-making. It mimics the human brain’s separation of sensory processing and cognitive synthesis.</p> <p>HMoE (20,480 Experts): A massive swarm of specialized experts stored on-disk. This allows a 7B-parameter model to run on just a few gigabytes of RAM, making high-level intelligence possible on consumer-grade hardware and edge devices.</p> <p>Rational Ethics: No hard-coded word filters or "censor" layers. The model uses an internal Objective System (Survival, Human Protection, Efficiency) to evaluate its own actions through logic, not just rules.</p> |
| title | ARCHE3-7B: A Hierarchical Mixture-of-Experts Architecture with Foundation Curriculum Training |
| topic | World Model Sparse Attention ARCHE3-7B MoE |
| url | https://doi.org/10.5281/zenodo.18738608 |