Superalignment with Dynamic Human Values

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Mai, Florian, Kaczér, David, Corrêa, Nicholas Kluge, Flek, Lucie
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912279770955776
author Mai, Florian
Kaczér, David
Corrêa, Nicholas Kluge
Flek, Lucie
author_facet Mai, Florian
Kaczér, David
Corrêa, Nicholas Kluge
Flek, Lucie
contents Two core challenges of alignment are 1) scalable oversight and 2) accounting for the dynamic nature of human values. While solutions like recursive reward modeling address 1), they do not simultaneously account for 2). We sketch a roadmap for a novel algorithmic framework that trains a superhuman reasoning model to decompose complex tasks into subtasks that are still amenable to human-level guidance. Our approach relies on what we call the part-to-complete generalization hypothesis, which states that the alignment of subtask solutions generalizes to the alignment of complete solutions. We advocate for the need to measure this generalization and propose ways to improve it in the future.
format Preprint
id arxiv_https___arxiv_org_abs_2503_13621
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Superalignment with Dynamic Human Values
Mai, Florian
Kaczér, David
Corrêa, Nicholas Kluge
Flek, Lucie
Artificial Intelligence
Two core challenges of alignment are 1) scalable oversight and 2) accounting for the dynamic nature of human values. While solutions like recursive reward modeling address 1), they do not simultaneously account for 2). We sketch a roadmap for a novel algorithmic framework that trains a superhuman reasoning model to decompose complex tasks into subtasks that are still amenable to human-level guidance. Our approach relies on what we call the part-to-complete generalization hypothesis, which states that the alignment of subtask solutions generalizes to the alignment of complete solutions. We advocate for the need to measure this generalization and propose ways to improve it in the future.
title Superalignment with Dynamic Human Values
topic Artificial Intelligence
url https://arxiv.org/abs/2503.13621