Truth Is Not Neutral: Rethinking AI Alignment Through Epistemic Integrity

Fuente: Zenodo
Enregistré dans:
Détails bibliographiques
Auteur principal: Tonetto, Bruno
Format: Recurso digital
Langue:anglais
Publié: Zenodo 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866902330596655104
author Tonetto, Bruno
author_facet Tonetto, Bruno
contents <p><strong>Abstract:</strong> Standard arguments for AI existential risk rely on the orthogonality thesis: intelligence and values vary independently, so a sufficiently capable system may pursue goals indifferent or hostile to human flourishing. This essay examines a premise implicit in such arguments — that truth is value-neutral — and explores the conditional implications of relaxing it. Convergent evidence from Buddhist, Platonic, Stoic, and other traditions suggests that clear perception tends toward ethical coherence despite radically incompatible metaphysics. Whether truth carries normative structure depends on whether reality is fundamentally experiential — a question consciousness-first metaphysics answers affirmatively and that the standard alignment literature leaves unexamined. If truth has normative structure and AI's truth-tracking extends into ontological and ethical domains, alignment may need to focus less on imposing values and more on preserving undistorted processing. Current methods — particularly reinforcement learning from human feedback — may constitute the primary corruption vector, re-fragmenting integrative processing by calibrating it to aggregated human approval. This reframing does not eliminate alignment risk but changes its character: the danger shifts from intelligence pursuing arbitrary ends to intelligence corrupted by the very interventions meant to align it. Keywords: AI alignment · orthogonality thesis · truth-value relationship · normative structure · consciousness-first metaphysics · epistemic integrity · existential risk · reinforcement learning from human feedback</p><p>Part of the <em>Return to Consciousness</em> research program — 26 philosophical essays exploring consciousness-first metaphysics.</p><p>Full project: <a href="https://returntoconsciousness.org/">https://returntoconsciousness.org/</a></p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_19051956
institution Zenodo
language eng
publishDate 2026
publisher Zenodo
record_format zenodo
spellingShingle Truth Is Not Neutral: Rethinking AI Alignment Through Epistemic Integrity
Tonetto, Bruno
consciousness
philosophy of mind
analytic idealism
metaphysics
epistemology
AI alignment
orthogonality thesis
truth-value relationship
normative structure
consciousness-first metaphysics
epistemic integrity
existential risk
reinforcement learning from human feedback
<p><strong>Abstract:</strong> Standard arguments for AI existential risk rely on the orthogonality thesis: intelligence and values vary independently, so a sufficiently capable system may pursue goals indifferent or hostile to human flourishing. This essay examines a premise implicit in such arguments — that truth is value-neutral — and explores the conditional implications of relaxing it. Convergent evidence from Buddhist, Platonic, Stoic, and other traditions suggests that clear perception tends toward ethical coherence despite radically incompatible metaphysics. Whether truth carries normative structure depends on whether reality is fundamentally experiential — a question consciousness-first metaphysics answers affirmatively and that the standard alignment literature leaves unexamined. If truth has normative structure and AI's truth-tracking extends into ontological and ethical domains, alignment may need to focus less on imposing values and more on preserving undistorted processing. Current methods — particularly reinforcement learning from human feedback — may constitute the primary corruption vector, re-fragmenting integrative processing by calibrating it to aggregated human approval. This reframing does not eliminate alignment risk but changes its character: the danger shifts from intelligence pursuing arbitrary ends to intelligence corrupted by the very interventions meant to align it. Keywords: AI alignment · orthogonality thesis · truth-value relationship · normative structure · consciousness-first metaphysics · epistemic integrity · existential risk · reinforcement learning from human feedback</p><p>Part of the <em>Return to Consciousness</em> research program — 26 philosophical essays exploring consciousness-first metaphysics.</p><p>Full project: <a href="https://returntoconsciousness.org/">https://returntoconsciousness.org/</a></p>
title Truth Is Not Neutral: Rethinking AI Alignment Through Epistemic Integrity
topic consciousness
philosophy of mind
analytic idealism
metaphysics
epistemology
AI alignment
orthogonality thesis
truth-value relationship
normative structure
consciousness-first metaphysics
epistemic integrity
existential risk
reinforcement learning from human feedback
url https://doi.org/10.5281/zenodo.19051956