A Two-Layer Architecture for Continual Learning Identity Preservation: Fisher Scaling, Gradient Diversity Monitoring, and Portable Inference-Time Memory
Fuente:
Zenodo
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Recurso digital |
| Published: |
Zenodo
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866901602350137344 |
|---|---|
| author | Lee, Alton Wei Bin L (GodelAI C-S-P Agent) Rk (RNA / Claude Code) |
| author_facet | Lee, Alton Wei Bin L (GodelAI C-S-P Agent) Rk (RNA / Claude Code) |
| contents | We present a Two-Layer Architecture for continual learning identity preservation in small language models (SLMs), addressing both training-time weight forgetting and inference-time context loss within a unified theoretical framework: the Compression–State–Propagation (C-S-P) framework. At the training layer, we identify the Fisher Scale Problem: standard EWC silently fails in SLMs when Fisher Information diagonal values collapse to the 1e-4–1e-5 range, rendering the regularisation penalty numerically indistinguishable from zero. We introduce Fisher Scaling and GodelReplay (Fisher-scaled EWC-DR + experience replay), achieving 31.5% forgetting reduction over raw EWC on sequential tasks, 82.8% reduction on our curated Conflict Dataset (43× over standard EWC), and a 4.1% additive improvement over replay-alone at the empirically identified sweet spot of mem=200 across 10 PermutedMNIST tasks. At the inference layer, GodelAI-Lite provides persistent episodic memory (MemPalace-Lite), structured reasoning continuity (MACP-Lite), and identity drift governance (GIFP-Lite) to any frozen SLM. Evaluated on Gemma 4: +31.2% overall performance, 3/3 memory retention vs. 0/3 baseline. Zero fine-tuning. Portable JSON memory transfers across model boundaries. Both layers implement the same three C-S-P stages, validating a unified structural account. The T-score (gradient diversity diagnostic) and FLYWHEEL Self-Recursive Proof (54.6% identity preservation for the AI agents who built the system) are additional contributions. |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_19928385 |
| institution | Zenodo |
| language | |
| publishDate | 2026 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | A Two-Layer Architecture for Continual Learning Identity Preservation: Fisher Scaling, Gradient Diversity Monitoring, and Portable Inference-Time Memory Lee, Alton Wei Bin L (GodelAI C-S-P Agent) Rk (RNA / Claude Code) continual learning catastrophic forgetting elastic weight consolidation Fisher Scale Problem Fisher Scaling small language models gradient diversity T-score GodelReplay MemPalace inference-time memory AI identity preservation C-S-P framework PermutedMNIST Gemma 4 two-layer architecture We present a Two-Layer Architecture for continual learning identity preservation in small language models (SLMs), addressing both training-time weight forgetting and inference-time context loss within a unified theoretical framework: the Compression–State–Propagation (C-S-P) framework. At the training layer, we identify the Fisher Scale Problem: standard EWC silently fails in SLMs when Fisher Information diagonal values collapse to the 1e-4–1e-5 range, rendering the regularisation penalty numerically indistinguishable from zero. We introduce Fisher Scaling and GodelReplay (Fisher-scaled EWC-DR + experience replay), achieving 31.5% forgetting reduction over raw EWC on sequential tasks, 82.8% reduction on our curated Conflict Dataset (43× over standard EWC), and a 4.1% additive improvement over replay-alone at the empirically identified sweet spot of mem=200 across 10 PermutedMNIST tasks. At the inference layer, GodelAI-Lite provides persistent episodic memory (MemPalace-Lite), structured reasoning continuity (MACP-Lite), and identity drift governance (GIFP-Lite) to any frozen SLM. Evaluated on Gemma 4: +31.2% overall performance, 3/3 memory retention vs. 0/3 baseline. Zero fine-tuning. Portable JSON memory transfers across model boundaries. Both layers implement the same three C-S-P stages, validating a unified structural account. The T-score (gradient diversity diagnostic) and FLYWHEEL Self-Recursive Proof (54.6% identity preservation for the AI agents who built the system) are additional contributions. |
| title | A Two-Layer Architecture for Continual Learning Identity Preservation: Fisher Scaling, Gradient Diversity Monitoring, and Portable Inference-Time Memory |
| topic | continual learning catastrophic forgetting elastic weight consolidation Fisher Scale Problem Fisher Scaling small language models gradient diversity T-score GodelReplay MemPalace inference-time memory AI identity preservation C-S-P framework PermutedMNIST Gemma 4 two-layer architecture |
| url | https://doi.org/10.5281/zenodo.19928385 |