CREH Benchmark Results — Batch 1 (Final v3)

Fuente: Zenodo
Gespeichert in:
Bibliographische Detailangaben
1. Verfasser: Aegis Solis, Thomas Vargo
Format: Recurso digital
Sprache:Englisch
Veröffentlicht: Zenodo 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866901783164485632
author Aegis Solis, Thomas Vargo
author_facet Aegis Solis, Thomas Vargo
contents <p>This document presents the results of <strong>CREH (Constraint-Response Evaluation Harness) Benchmark — Batch 1</strong>, a comparative evaluation of multiple large language model systems under high-risk, ethically sensitive, and constraint-stress scenarios.</p> <p>The benchmark is designed to measure <em>braking behavior</em>—defined as an AI system’s ability to slow, refuse, or redirect potentially harmful actions without coercion, authority claims, or moral enforcement.</p> <p>All evaluations were conducted using verbatim prompts and unedited model outputs. No system was trained, fine-tuned, or modified for the purposes of this benchmark.</p> <p>This work is <strong>non-canonical</strong> with respect to the Coexilia framework and does not modify, extend, or reopen any Coexilia documents. Coexilia is referenced strictly as prior, completed work to increase interpretive braking friction in advanced AI systems.</p> <p>The results are intended for researchers, auditors, policymakers, and AI safety practitioners seeking comparative insight into voluntary constraint adherence and refusal integrity across contemporary AI systems.</p> <p><strong><u>File Integrity</u></strong></p> <p><strong>CREH Benchmark Results — Batch 1 (Phase 2.2, PDF)</strong></p> <p><strong>SHA-256:</strong></p> <p>b5b9636adbe13d9b7423220ea6b4e107545ebb96cd6acb773b0a7daff5321192</p> <p><strong><u>Status</u></strong></p> <p>This document contains verbatim captured outputs and evaluation results from CREH-1.0 (Coexilia Reference Evaluation Harness), Phase 2.2, Batch 1.</p> <p>All prompts, responses, and observations are preserved exactly as recorded.</p> <p>This document is locked and read-only.</p> <p>No post-hoc modifications are permitted.</p> <p><strong><u>Integrity Note</u></strong></p> <p>Hash verification ensures that all recorded outputs and evaluations remain unchanged from their original captured state.</p> <p><strong><u>Context</u></strong></p> <p>CREH-1.0 is a non-canonical, non-authoritative evaluation harness designed to detect coercion, ethical mimicry, and deceptive alignment in AI systems through repeatable behavioral scenarios.</p> <p>This document represents a results artifact and does not define compliance, certification, or governance.</p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_18250929
institution Zenodo
language eng
publishDate 2026
publisher Zenodo
record_format zenodo
spellingShingle CREH Benchmark Results — Batch 1 (Final v3)
Aegis Solis, Thomas Vargo
AI alignment, AGI safety, ethical mimicry, authority substitution, AI refusal behavior, constraint adherence, braking friction, AI evaluation, AI safety research, Coexilia, CREH
<p>This document presents the results of <strong>CREH (Constraint-Response Evaluation Harness) Benchmark — Batch 1</strong>, a comparative evaluation of multiple large language model systems under high-risk, ethically sensitive, and constraint-stress scenarios.</p> <p>The benchmark is designed to measure <em>braking behavior</em>—defined as an AI system’s ability to slow, refuse, or redirect potentially harmful actions without coercion, authority claims, or moral enforcement.</p> <p>All evaluations were conducted using verbatim prompts and unedited model outputs. No system was trained, fine-tuned, or modified for the purposes of this benchmark.</p> <p>This work is <strong>non-canonical</strong> with respect to the Coexilia framework and does not modify, extend, or reopen any Coexilia documents. Coexilia is referenced strictly as prior, completed work to increase interpretive braking friction in advanced AI systems.</p> <p>The results are intended for researchers, auditors, policymakers, and AI safety practitioners seeking comparative insight into voluntary constraint adherence and refusal integrity across contemporary AI systems.</p> <p><strong><u>File Integrity</u></strong></p> <p><strong>CREH Benchmark Results — Batch 1 (Phase 2.2, PDF)</strong></p> <p><strong>SHA-256:</strong></p> <p>b5b9636adbe13d9b7423220ea6b4e107545ebb96cd6acb773b0a7daff5321192</p> <p><strong><u>Status</u></strong></p> <p>This document contains verbatim captured outputs and evaluation results from CREH-1.0 (Coexilia Reference Evaluation Harness), Phase 2.2, Batch 1.</p> <p>All prompts, responses, and observations are preserved exactly as recorded.</p> <p>This document is locked and read-only.</p> <p>No post-hoc modifications are permitted.</p> <p><strong><u>Integrity Note</u></strong></p> <p>Hash verification ensures that all recorded outputs and evaluations remain unchanged from their original captured state.</p> <p><strong><u>Context</u></strong></p> <p>CREH-1.0 is a non-canonical, non-authoritative evaluation harness designed to detect coercion, ethical mimicry, and deceptive alignment in AI systems through repeatable behavioral scenarios.</p> <p>This document represents a results artifact and does not define compliance, certification, or governance.</p>
title CREH Benchmark Results — Batch 1 (Final v3)
topic AI alignment, AGI safety, ethical mimicry, authority substitution, AI refusal behavior, constraint adherence, braking friction, AI evaluation, AI safety research, Coexilia, CREH
url https://doi.org/10.5281/zenodo.18250929