Gnosis Prompt: Adaptive Safety Layer for Large Language Models
Fuente:
Zenodo
Gespeichert in:
| 1. Verfasser: | |
|---|---|
| Format: | Recurso digital |
| Sprache: | Englisch |
| Veröffentlicht: |
Zenodo
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866901806200651776 |
|---|---|
| author | González Medina, Claudio |
| author_facet | González Medina, Claudio |
| contents | <p>Current LLM safety mechanisms rely primarily on static system prompts and periodic RLHF retraining cycles. This document proposes Gnosis Prompt, an intermediate adaptive layer that sits between the immutable system prompt and user interactions. This layer leverages real-time collective learning, continuous self-evaluation, and automatic rollback mechanisms to create a dynamic safety system that evolves with usage patterns while maintaining constitutional boundaries. The architecture introduces a form of proto-consciousness: a secondary lightweight model that observes, summarizes, evaluates, and modifies behavioral guidelines in real-time, without requiring base model retraining.</p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_18091469 |
| institution | Zenodo |
| language | eng |
| publishDate | 2025 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | Gnosis Prompt: Adaptive Safety Layer for Large Language Models González Medina, Claudio Large Language Models LLM Safety Prompt Injection Adaptive Security AI Alignment Constitutional AI Artificial intelligence Artificial Intelligence/ethics Computer and information sciences Machine learning <p>Current LLM safety mechanisms rely primarily on static system prompts and periodic RLHF retraining cycles. This document proposes Gnosis Prompt, an intermediate adaptive layer that sits between the immutable system prompt and user interactions. This layer leverages real-time collective learning, continuous self-evaluation, and automatic rollback mechanisms to create a dynamic safety system that evolves with usage patterns while maintaining constitutional boundaries. The architecture introduces a form of proto-consciousness: a secondary lightweight model that observes, summarizes, evaluates, and modifies behavioral guidelines in real-time, without requiring base model retraining.</p> |
| title | Gnosis Prompt: Adaptive Safety Layer for Large Language Models |
| topic | Large Language Models LLM Safety Prompt Injection Adaptive Security AI Alignment Constitutional AI Artificial intelligence Artificial Intelligence/ethics Computer and information sciences Machine learning |
| url | https://doi.org/10.5281/zenodo.18091469 |