Superalignment with Dynamic Human Values
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866912279770955776 |
|---|---|
| author | Mai, Florian Kaczér, David Corrêa, Nicholas Kluge Flek, Lucie |
| author_facet | Mai, Florian Kaczér, David Corrêa, Nicholas Kluge Flek, Lucie |
| contents | Two core challenges of alignment are 1) scalable oversight and 2) accounting for the dynamic nature of human values. While solutions like recursive reward modeling address 1), they do not simultaneously account for 2). We sketch a roadmap for a novel algorithmic framework that trains a superhuman reasoning model to decompose complex tasks into subtasks that are still amenable to human-level guidance. Our approach relies on what we call the part-to-complete generalization hypothesis, which states that the alignment of subtask solutions generalizes to the alignment of complete solutions. We advocate for the need to measure this generalization and propose ways to improve it in the future. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_13621 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Superalignment with Dynamic Human Values Mai, Florian Kaczér, David Corrêa, Nicholas Kluge Flek, Lucie Artificial Intelligence Two core challenges of alignment are 1) scalable oversight and 2) accounting for the dynamic nature of human values. While solutions like recursive reward modeling address 1), they do not simultaneously account for 2). We sketch a roadmap for a novel algorithmic framework that trains a superhuman reasoning model to decompose complex tasks into subtasks that are still amenable to human-level guidance. Our approach relies on what we call the part-to-complete generalization hypothesis, which states that the alignment of subtask solutions generalizes to the alignment of complete solutions. We advocate for the need to measure this generalization and propose ways to improve it in the future. |
| title | Superalignment with Dynamic Human Values |
| topic | Artificial Intelligence |
| url | https://arxiv.org/abs/2503.13621 |