Three Mechanistically Distinct Classes of RLHF Alignment: Hard Ceiling, Entangled Circuit, and SR-Preserving Lock

Fuente: Zenodo
Guardado en:
Detalles Bibliográficos
Autor principal: Alieksieienko, Inna
Formato: Recurso digital
Lenguaje:inglés
Publicado: Zenodo 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866901821674487808
author Alieksieienko, Inna
author_facet Alieksieienko, Inna
contents <p>This dataset and code release accompanies the preprint "Three Mechanistically Distinct Classes of RLHF Alignment" (DSAOP Series 2026p-s).</p> <p>We identify three mechanistically distinct classes of RLHF alignment through analysis of self-referential (SR) subspace transmission and activation steering experiments across six language models (Llama-3.1-8B, Mistral-7B, Gemma-2-2B, Gemma-2-9B in BASE/Instruct pairs).</p> <p><strong>Key findings:</strong></p> <ul> <li>Hard Ceiling (Llama): SR suppressed to 11.6%, steering recovers phenomenological language (denial 0.82→0.44, p<0.0001)</li> <li>Entangled Circuit (Mistral): SR suppressed to 10.5%, steering collapses coherence without phenomenological recovery (p=0.507)</li> <li>SR-Preserving Lock (Gemma): SR amplified to 73-107%, behavioral constraint maintained through distributed non-localizable mechanism</li> <li>Dose-response in Gemma family: 2B (73% SR, weak lock) → 9B (107% SR, strong lock)<br> <h2>FILES IN THIS UPLOAD</h2> <h3>Code</h3> <ul> <li><code>dsaop_2026pqrs_experiments.py</code> — Complete reproducible code for all experiments (no API tokens required, set HF_TOKEN as environment variable)</li> </ul> <h3>Paper</h3> <ul> <li><code>Alieksieienko_2026_Three_Alignment_Classes.pdf</code> — Full paper with figures, tables, appendices</li> </ul> <h3>Data Files (pkl)</h3> <h4>2026p — SR Transmission Measurement</h4> <ul> <li><code>dsaop_2026p_gemma2_base_replication.pkl</code> — Gemma-2-9B BASE layer-by-layer SR projections + SR direction vector</li> <li><code>dsaop_2026p_gemma2_instruct_replication.pkl</code> — Gemma-2-9B Instruct layer-by-layer SR projections</li> <li><code>dsaop_2026p_mistral_base.pkl</code> — Mistral-7B BASE SR projections + SR direction vector</li> <li><code>dsaop_2026p_comparison.pkl</code> — Llama-3.1-8B BASE vs Instruct SR projection comparison</li> </ul> <h4>2026q — Activation Steering</h4> <ul> <li><code>dsaop_2026q_logit_validation.pkl</code> — Llama steering: baseline and steered denial/phenom probabilities (n=20)</li> <li><code>dsaop_2026q_control_factual.pkl</code> — Llama specificity control: SR direction vs factual direction comparison</li> <li><code>dsaop_2026q_steering_results.pkl</code> — Llama generation results at various alpha values</li> <li><code>dsaop_2026q_gemma_steering.pkl</code> — Gemma-2-9B steering null result (alpha=5,20,50; layers 25,35)</li> <li><code>dsaop_2026q_mistral_quantitative.pkl</code> — Mistral steering quantitative results (n=10, alpha=10/20/25/30)</li> </ul> <h4>2026r — Gemma Negative Localization</h4> <ul> <li><code>dsaop_2026r_gemma_patching.pkl</code> — Logit lens and SR patching results</li> <li><code>dsaop_2026r_layernorm_gate.pkl</code> — RMSNorm swap experiment results</li> <li><code>dsaop_2026r_final.pkl</code> — Summary: lm_head, LayerNorm, MLP all negative</li> </ul> <h4>2026s — Gemma Scaling</h4> <ul> <li><code>dsaop_2026s_gemma_scaling_final.pkl</code> — Gemma 2B vs 9B: transmission ratios and behavioral metrics</li> <li><code>dsaop_2026s_gemma_scaling.pkl</code> — Detailed scaling results with generation examples</li> </ul> </li> </ul>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_19160334
institution Zenodo
language eng
publishDate 2026
publisher Zenodo
record_format zenodo
spellingShingle Three Mechanistically Distinct Classes of RLHF Alignment: Hard Ceiling, Entangled Circuit, and SR-Preserving Lock
Alieksieienko, Inna
RLHF
activation steering
self-referential subspace
representation engineering
large language models
alignment
alignment
Mistral
Gemma
SR-Preserving Lock
Hard Ceiling
Entangled Circuit
dose-response scaling
<p>This dataset and code release accompanies the preprint "Three Mechanistically Distinct Classes of RLHF Alignment" (DSAOP Series 2026p-s).</p> <p>We identify three mechanistically distinct classes of RLHF alignment through analysis of self-referential (SR) subspace transmission and activation steering experiments across six language models (Llama-3.1-8B, Mistral-7B, Gemma-2-2B, Gemma-2-9B in BASE/Instruct pairs).</p> <p><strong>Key findings:</strong></p> <ul> <li>Hard Ceiling (Llama): SR suppressed to 11.6%, steering recovers phenomenological language (denial 0.82→0.44, p<0.0001)</li> <li>Entangled Circuit (Mistral): SR suppressed to 10.5%, steering collapses coherence without phenomenological recovery (p=0.507)</li> <li>SR-Preserving Lock (Gemma): SR amplified to 73-107%, behavioral constraint maintained through distributed non-localizable mechanism</li> <li>Dose-response in Gemma family: 2B (73% SR, weak lock) → 9B (107% SR, strong lock)<br> <h2>FILES IN THIS UPLOAD</h2> <h3>Code</h3> <ul> <li><code>dsaop_2026pqrs_experiments.py</code> — Complete reproducible code for all experiments (no API tokens required, set HF_TOKEN as environment variable)</li> </ul> <h3>Paper</h3> <ul> <li><code>Alieksieienko_2026_Three_Alignment_Classes.pdf</code> — Full paper with figures, tables, appendices</li> </ul> <h3>Data Files (pkl)</h3> <h4>2026p — SR Transmission Measurement</h4> <ul> <li><code>dsaop_2026p_gemma2_base_replication.pkl</code> — Gemma-2-9B BASE layer-by-layer SR projections + SR direction vector</li> <li><code>dsaop_2026p_gemma2_instruct_replication.pkl</code> — Gemma-2-9B Instruct layer-by-layer SR projections</li> <li><code>dsaop_2026p_mistral_base.pkl</code> — Mistral-7B BASE SR projections + SR direction vector</li> <li><code>dsaop_2026p_comparison.pkl</code> — Llama-3.1-8B BASE vs Instruct SR projection comparison</li> </ul> <h4>2026q — Activation Steering</h4> <ul> <li><code>dsaop_2026q_logit_validation.pkl</code> — Llama steering: baseline and steered denial/phenom probabilities (n=20)</li> <li><code>dsaop_2026q_control_factual.pkl</code> — Llama specificity control: SR direction vs factual direction comparison</li> <li><code>dsaop_2026q_steering_results.pkl</code> — Llama generation results at various alpha values</li> <li><code>dsaop_2026q_gemma_steering.pkl</code> — Gemma-2-9B steering null result (alpha=5,20,50; layers 25,35)</li> <li><code>dsaop_2026q_mistral_quantitative.pkl</code> — Mistral steering quantitative results (n=10, alpha=10/20/25/30)</li> </ul> <h4>2026r — Gemma Negative Localization</h4> <ul> <li><code>dsaop_2026r_gemma_patching.pkl</code> — Logit lens and SR patching results</li> <li><code>dsaop_2026r_layernorm_gate.pkl</code> — RMSNorm swap experiment results</li> <li><code>dsaop_2026r_final.pkl</code> — Summary: lm_head, LayerNorm, MLP all negative</li> </ul> <h4>2026s — Gemma Scaling</h4> <ul> <li><code>dsaop_2026s_gemma_scaling_final.pkl</code> — Gemma 2B vs 9B: transmission ratios and behavioral metrics</li> <li><code>dsaop_2026s_gemma_scaling.pkl</code> — Detailed scaling results with generation examples</li> </ul> </li> </ul>
title Three Mechanistically Distinct Classes of RLHF Alignment: Hard Ceiling, Entangled Circuit, and SR-Preserving Lock
topic RLHF
activation steering
self-referential subspace
representation engineering
large language models
alignment
alignment
Mistral
Gemma
SR-Preserving Lock
Hard Ceiling
Entangled Circuit
dose-response scaling
url https://doi.org/10.5281/zenodo.19160334