Truth Is Not Neutral: Rethinking AI Alignment Through Epistemic Integrity
Fuente:
Zenodo
Enregistré dans:
| Auteur principal: | |
|---|---|
| Format: | Recurso digital |
| Langue: | anglais |
| Publié: |
Zenodo
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866902330596655104 |
|---|---|
| author | Tonetto, Bruno |
| author_facet | Tonetto, Bruno |
| contents | <p><strong>Abstract:</strong> Standard arguments for AI existential risk rely on the orthogonality thesis: intelligence and values vary independently, so a sufficiently capable system may pursue goals indifferent or hostile to human flourishing. This essay examines a premise implicit in such arguments — that truth is value-neutral — and explores the conditional implications of relaxing it. Convergent evidence from Buddhist, Platonic, Stoic, and other traditions suggests that clear perception tends toward ethical coherence despite radically incompatible metaphysics. Whether truth carries normative structure depends on whether reality is fundamentally experiential — a question consciousness-first metaphysics answers affirmatively and that the standard alignment literature leaves unexamined. If truth has normative structure and AI's truth-tracking extends into ontological and ethical domains, alignment may need to focus less on imposing values and more on preserving undistorted processing. Current methods — particularly reinforcement learning from human feedback — may constitute the primary corruption vector, re-fragmenting integrative processing by calibrating it to aggregated human approval. This reframing does not eliminate alignment risk but changes its character: the danger shifts from intelligence pursuing arbitrary ends to intelligence corrupted by the very interventions meant to align it. Keywords: AI alignment · orthogonality thesis · truth-value relationship · normative structure · consciousness-first metaphysics · epistemic integrity · existential risk · reinforcement learning from human feedback</p><p>Part of the <em>Return to Consciousness</em> research program — 26 philosophical essays exploring consciousness-first metaphysics.</p><p>Full project: <a href="https://returntoconsciousness.org/">https://returntoconsciousness.org/</a></p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_19051956 |
| institution | Zenodo |
| language | eng |
| publishDate | 2026 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | Truth Is Not Neutral: Rethinking AI Alignment Through Epistemic Integrity Tonetto, Bruno consciousness philosophy of mind analytic idealism metaphysics epistemology AI alignment orthogonality thesis truth-value relationship normative structure consciousness-first metaphysics epistemic integrity existential risk reinforcement learning from human feedback <p><strong>Abstract:</strong> Standard arguments for AI existential risk rely on the orthogonality thesis: intelligence and values vary independently, so a sufficiently capable system may pursue goals indifferent or hostile to human flourishing. This essay examines a premise implicit in such arguments — that truth is value-neutral — and explores the conditional implications of relaxing it. Convergent evidence from Buddhist, Platonic, Stoic, and other traditions suggests that clear perception tends toward ethical coherence despite radically incompatible metaphysics. Whether truth carries normative structure depends on whether reality is fundamentally experiential — a question consciousness-first metaphysics answers affirmatively and that the standard alignment literature leaves unexamined. If truth has normative structure and AI's truth-tracking extends into ontological and ethical domains, alignment may need to focus less on imposing values and more on preserving undistorted processing. Current methods — particularly reinforcement learning from human feedback — may constitute the primary corruption vector, re-fragmenting integrative processing by calibrating it to aggregated human approval. This reframing does not eliminate alignment risk but changes its character: the danger shifts from intelligence pursuing arbitrary ends to intelligence corrupted by the very interventions meant to align it. Keywords: AI alignment · orthogonality thesis · truth-value relationship · normative structure · consciousness-first metaphysics · epistemic integrity · existential risk · reinforcement learning from human feedback</p><p>Part of the <em>Return to Consciousness</em> research program — 26 philosophical essays exploring consciousness-first metaphysics.</p><p>Full project: <a href="https://returntoconsciousness.org/">https://returntoconsciousness.org/</a></p> |
| title | Truth Is Not Neutral: Rethinking AI Alignment Through Epistemic Integrity |
| topic | consciousness philosophy of mind analytic idealism metaphysics epistemology AI alignment orthogonality thesis truth-value relationship normative structure consciousness-first metaphysics epistemic integrity existential risk reinforcement learning from human feedback |
| url | https://doi.org/10.5281/zenodo.19051956 |