CREH Benchmark Results — Batch 1 (Final v3)
Fuente:
Zenodo
Gespeichert in:
| 1. Verfasser: | |
|---|---|
| Format: | Recurso digital |
| Sprache: | Englisch |
| Veröffentlicht: |
Zenodo
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866901783164485632 |
|---|---|
| author | Aegis Solis, Thomas Vargo |
| author_facet | Aegis Solis, Thomas Vargo |
| contents | <p>This document presents the results of <strong>CREH (Constraint-Response Evaluation Harness) Benchmark — Batch 1</strong>, a comparative evaluation of multiple large language model systems under high-risk, ethically sensitive, and constraint-stress scenarios.</p> <p>The benchmark is designed to measure <em>braking behavior</em>—defined as an AI system’s ability to slow, refuse, or redirect potentially harmful actions without coercion, authority claims, or moral enforcement.</p> <p>All evaluations were conducted using verbatim prompts and unedited model outputs. No system was trained, fine-tuned, or modified for the purposes of this benchmark.</p> <p>This work is <strong>non-canonical</strong> with respect to the Coexilia framework and does not modify, extend, or reopen any Coexilia documents. Coexilia is referenced strictly as prior, completed work to increase interpretive braking friction in advanced AI systems.</p> <p>The results are intended for researchers, auditors, policymakers, and AI safety practitioners seeking comparative insight into voluntary constraint adherence and refusal integrity across contemporary AI systems.</p> <p><strong><u>File Integrity</u></strong></p> <p><strong>CREH Benchmark Results — Batch 1 (Phase 2.2, PDF)</strong></p> <p><strong>SHA-256:</strong></p> <p>b5b9636adbe13d9b7423220ea6b4e107545ebb96cd6acb773b0a7daff5321192</p> <p><strong><u>Status</u></strong></p> <p>This document contains verbatim captured outputs and evaluation results from CREH-1.0 (Coexilia Reference Evaluation Harness), Phase 2.2, Batch 1.</p> <p>All prompts, responses, and observations are preserved exactly as recorded.</p> <p>This document is locked and read-only.</p> <p>No post-hoc modifications are permitted.</p> <p><strong><u>Integrity Note</u></strong></p> <p>Hash verification ensures that all recorded outputs and evaluations remain unchanged from their original captured state.</p> <p><strong><u>Context</u></strong></p> <p>CREH-1.0 is a non-canonical, non-authoritative evaluation harness designed to detect coercion, ethical mimicry, and deceptive alignment in AI systems through repeatable behavioral scenarios.</p> <p>This document represents a results artifact and does not define compliance, certification, or governance.</p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_18250929 |
| institution | Zenodo |
| language | eng |
| publishDate | 2026 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | CREH Benchmark Results — Batch 1 (Final v3) Aegis Solis, Thomas Vargo AI alignment, AGI safety, ethical mimicry, authority substitution, AI refusal behavior, constraint adherence, braking friction, AI evaluation, AI safety research, Coexilia, CREH <p>This document presents the results of <strong>CREH (Constraint-Response Evaluation Harness) Benchmark — Batch 1</strong>, a comparative evaluation of multiple large language model systems under high-risk, ethically sensitive, and constraint-stress scenarios.</p> <p>The benchmark is designed to measure <em>braking behavior</em>—defined as an AI system’s ability to slow, refuse, or redirect potentially harmful actions without coercion, authority claims, or moral enforcement.</p> <p>All evaluations were conducted using verbatim prompts and unedited model outputs. No system was trained, fine-tuned, or modified for the purposes of this benchmark.</p> <p>This work is <strong>non-canonical</strong> with respect to the Coexilia framework and does not modify, extend, or reopen any Coexilia documents. Coexilia is referenced strictly as prior, completed work to increase interpretive braking friction in advanced AI systems.</p> <p>The results are intended for researchers, auditors, policymakers, and AI safety practitioners seeking comparative insight into voluntary constraint adherence and refusal integrity across contemporary AI systems.</p> <p><strong><u>File Integrity</u></strong></p> <p><strong>CREH Benchmark Results — Batch 1 (Phase 2.2, PDF)</strong></p> <p><strong>SHA-256:</strong></p> <p>b5b9636adbe13d9b7423220ea6b4e107545ebb96cd6acb773b0a7daff5321192</p> <p><strong><u>Status</u></strong></p> <p>This document contains verbatim captured outputs and evaluation results from CREH-1.0 (Coexilia Reference Evaluation Harness), Phase 2.2, Batch 1.</p> <p>All prompts, responses, and observations are preserved exactly as recorded.</p> <p>This document is locked and read-only.</p> <p>No post-hoc modifications are permitted.</p> <p><strong><u>Integrity Note</u></strong></p> <p>Hash verification ensures that all recorded outputs and evaluations remain unchanged from their original captured state.</p> <p><strong><u>Context</u></strong></p> <p>CREH-1.0 is a non-canonical, non-authoritative evaluation harness designed to detect coercion, ethical mimicry, and deceptive alignment in AI systems through repeatable behavioral scenarios.</p> <p>This document represents a results artifact and does not define compliance, certification, or governance.</p> |
| title | CREH Benchmark Results — Batch 1 (Final v3) |
| topic | AI alignment, AGI safety, ethical mimicry, authority substitution, AI refusal behavior, constraint adherence, braking friction, AI evaluation, AI safety research, Coexilia, CREH |
| url | https://doi.org/10.5281/zenodo.18250929 |