| _version_ | 1866901231787573248 |
|---|---|
| author | Arleo, Carlos |
| author_facet | Arleo, Carlos |
| contents | <p>This paper addresses the limitations of current AI alignment methods like RLHF and Constitutional AI, which are insufficient for real-time, high-stakes governance. We introduce the Wisdom Forcing Function (WFF), a novel neurosymbolic and evolutionary architecture designed to solve this "governance gap." The WFF implements Frame-Based Principled Reasoning, combining a generative neural model with a deterministic symbolic verifier (the "Verified Dialectical Kernel") that enforces a machine-executable constitution. Its primary innovation is an evolutionary "immune system" that enables "metastable self-correction," allowing the system to diagnose, repair, and recover from constitutional violations in real-time. We present an analysis of the architecture and empirical evidence from system logs demonstrating its capacity for auditable, verifiable, and adaptive governance, positioning it as a proof-of-concept for a new class of "glass box" AI systems.</p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_17610427 |
| institution | Zenodo |
| language | |
| publishDate | 2025 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | Metastable Self-Correction: A Neurosymbolic and Evolutionary Architecture for Verifiable AI Governance. Arleo, Carlos <p>This paper addresses the limitations of current AI alignment methods like RLHF and Constitutional AI, which are insufficient for real-time, high-stakes governance. We introduce the Wisdom Forcing Function (WFF), a novel neurosymbolic and evolutionary architecture designed to solve this "governance gap." The WFF implements Frame-Based Principled Reasoning, combining a generative neural model with a deterministic symbolic verifier (the "Verified Dialectical Kernel") that enforces a machine-executable constitution. Its primary innovation is an evolutionary "immune system" that enables "metastable self-correction," allowing the system to diagnose, repair, and recover from constitutional violations in real-time. We present an analysis of the architecture and empirical evidence from system logs demonstrating its capacity for auditable, verifiable, and adaptive governance, positioning it as a proof-of-concept for a new class of "glass box" AI systems.</p> |
| title | Metastable Self-Correction: A Neurosymbolic and Evolutionary Architecture for Verifiable AI Governance. |
| url | https://doi.org/10.5281/zenodo.17610427 |