Gnosis Prompt: Adaptive Safety Layer for Large Language Models

Fuente: Zenodo
Gespeichert in:
Bibliographische Detailangaben
1. Verfasser: González Medina, Claudio
Format: Recurso digital
Sprache:Englisch
Veröffentlicht: Zenodo 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866901806200651776
author González Medina, Claudio
author_facet González Medina, Claudio
contents <p>Current LLM safety mechanisms rely primarily on static system prompts and periodic RLHF retraining cycles. This document proposes Gnosis Prompt, an intermediate adaptive layer that sits between the immutable system prompt and user interactions. This layer leverages real-time collective learning, continuous self-evaluation, and automatic rollback mechanisms to create a dynamic safety system that evolves with usage patterns while maintaining constitutional boundaries. The architecture introduces a form of proto-consciousness: a secondary lightweight model that observes, summarizes, evaluates, and modifies behavioral guidelines in real-time, without requiring base model retraining.</p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_18091469
institution Zenodo
language eng
publishDate 2025
publisher Zenodo
record_format zenodo
spellingShingle Gnosis Prompt: Adaptive Safety Layer for Large Language Models
González Medina, Claudio
Large Language Models
LLM Safety
Prompt Injection
Adaptive Security
AI Alignment
Constitutional AI
Artificial intelligence
Artificial Intelligence/ethics
Computer and information sciences
Machine learning
<p>Current LLM safety mechanisms rely primarily on static system prompts and periodic RLHF retraining cycles. This document proposes Gnosis Prompt, an intermediate adaptive layer that sits between the immutable system prompt and user interactions. This layer leverages real-time collective learning, continuous self-evaluation, and automatic rollback mechanisms to create a dynamic safety system that evolves with usage patterns while maintaining constitutional boundaries. The architecture introduces a form of proto-consciousness: a secondary lightweight model that observes, summarizes, evaluates, and modifies behavioral guidelines in real-time, without requiring base model retraining.</p>
title Gnosis Prompt: Adaptive Safety Layer for Large Language Models
topic Large Language Models
LLM Safety
Prompt Injection
Adaptive Security
AI Alignment
Constitutional AI
Artificial intelligence
Artificial Intelligence/ethics
Computer and information sciences
Machine learning
url https://doi.org/10.5281/zenodo.18091469