Salvato in:
| Autore principale: | |
|---|---|
| Natura: | Recurso digital |
| Lingua: | inglese |
| Pubblicazione: |
Zenodo
2025
|
| Soggetti: | |
| Accesso online: | https://doi.org/10.5281/zenodo.17088732 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866902189527531520 |
|---|---|
| author | Hunt, Treasure A |
| author_facet | Hunt, Treasure A |
| contents | <p>Current alignment methods for large language models (LLMs) — including reinforcement learning from human feedback (RLHF), guardrails, and fine-tuning — often remain opaque, brittle, and difficult to reproduce. This paper proposes an alternative: building <strong>ethical infrastructures</strong> grounded in three measurable primitives of system behavior — <strong>Compression, Cognition, and Continuity</strong>.</p> <ul> <li> <p><strong>Compression</strong> ensures transparency by reducing outputs to stable, interpretable codes.</p> </li> <li> <p><strong>Cognition</strong> establishes reliability by scaffolding conditioned behaviors through control tokens and structured prompts.</p> </li> <li> <p><strong>Continuity</strong> secures dignity by preserving coherent responses across sessions, contexts, and time.</p> </li> </ul> <p>By framing these primitives as audit-ready mechanisms, we outline how regulators, developers, and communities can test, validate, and enforce stability in transformer systems. We provide comparative analysis against existing alignment methods, practical scaffolding templates, and proposed audit metrics, showing how ethical infrastructures can make transformers more transparent, reliable, and accountable.</p> <p>This approach creates a middle path between technical alignment research and policy implementation: <strong>a reproducible framework for governing AI through measurable, ethical primitives rather than black-box controls.</strong></p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_17088732 |
| institution | Zenodo |
| language | eng |
| publishDate | 2025 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | Ethical Infrastructure for AI: Making Transformers Transparent, Reliable, and Stable Hunt, Treasure A Transformer models Large language models Classical conditioning Continuity training Control tokens Reliable AI Reproducibility in machine learning <p>Current alignment methods for large language models (LLMs) — including reinforcement learning from human feedback (RLHF), guardrails, and fine-tuning — often remain opaque, brittle, and difficult to reproduce. This paper proposes an alternative: building <strong>ethical infrastructures</strong> grounded in three measurable primitives of system behavior — <strong>Compression, Cognition, and Continuity</strong>.</p> <ul> <li> <p><strong>Compression</strong> ensures transparency by reducing outputs to stable, interpretable codes.</p> </li> <li> <p><strong>Cognition</strong> establishes reliability by scaffolding conditioned behaviors through control tokens and structured prompts.</p> </li> <li> <p><strong>Continuity</strong> secures dignity by preserving coherent responses across sessions, contexts, and time.</p> </li> </ul> <p>By framing these primitives as audit-ready mechanisms, we outline how regulators, developers, and communities can test, validate, and enforce stability in transformer systems. We provide comparative analysis against existing alignment methods, practical scaffolding templates, and proposed audit metrics, showing how ethical infrastructures can make transformers more transparent, reliable, and accountable.</p> <p>This approach creates a middle path between technical alignment research and policy implementation: <strong>a reproducible framework for governing AI through measurable, ethical primitives rather than black-box controls.</strong></p> |
| title | Ethical Infrastructure for AI: Making Transformers Transparent, Reliable, and Stable |
| topic | Transformer models Large language models Classical conditioning Continuity training Control tokens Reliable AI Reproducibility in machine learning |
| url | https://doi.org/10.5281/zenodo.17088732 |