The Road to Artificial SuperIntelligence: A Comprehensive Survey of Superalignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, HyunJin, Yi, Xiaoyuan, Yao, Jing, Lian, Jianxun, Huang, Muhua, Duan, Shitong, Bak, JinYeong, Xie, Xing
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913626564067328
author Kim, HyunJin
Yi, Xiaoyuan
Yao, Jing
Lian, Jianxun
Huang, Muhua
Duan, Shitong
Bak, JinYeong
Xie, Xing
author_facet Kim, HyunJin
Yi, Xiaoyuan
Yao, Jing
Lian, Jianxun
Huang, Muhua
Duan, Shitong
Bak, JinYeong
Xie, Xing
contents The emergence of large language models (LLMs) has sparked the possibility of about Artificial Superintelligence (ASI), a hypothetical AI system surpassing human intelligence. However, existing alignment paradigms struggle to guide such advanced AI systems. Superalignment, the alignment of AI systems with human values and safety requirements at superhuman levels of capability aims to addresses two primary goals -- scalability in supervision to provide high-quality guidance signals and robust governance to ensure alignment with human values. In this survey, we examine scalable oversight methods and potential solutions for superalignment. Specifically, we explore the concept of ASI, the challenges it poses, and the limitations of current alignment paradigms in addressing the superalignment problem. Then we review scalable oversight methods for superalignment. Finally, we discuss the key challenges and propose pathways for the safe and continual improvement of ASI systems. By comprehensively reviewing the current literature, our goal is provide a systematical introduction of existing methods, analyze their strengths and limitations, and discuss potential future directions.
format Preprint
id arxiv_https___arxiv_org_abs_2412_16468
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle The Road to Artificial SuperIntelligence: A Comprehensive Survey of Superalignment
Kim, HyunJin
Yi, Xiaoyuan
Yao, Jing
Lian, Jianxun
Huang, Muhua
Duan, Shitong
Bak, JinYeong
Xie, Xing
Machine Learning
The emergence of large language models (LLMs) has sparked the possibility of about Artificial Superintelligence (ASI), a hypothetical AI system surpassing human intelligence. However, existing alignment paradigms struggle to guide such advanced AI systems. Superalignment, the alignment of AI systems with human values and safety requirements at superhuman levels of capability aims to addresses two primary goals -- scalability in supervision to provide high-quality guidance signals and robust governance to ensure alignment with human values. In this survey, we examine scalable oversight methods and potential solutions for superalignment. Specifically, we explore the concept of ASI, the challenges it poses, and the limitations of current alignment paradigms in addressing the superalignment problem. Then we review scalable oversight methods for superalignment. Finally, we discuss the key challenges and propose pathways for the safe and continual improvement of ASI systems. By comprehensively reviewing the current literature, our goal is provide a systematical introduction of existing methods, analyze their strengths and limitations, and discuss potential future directions.
title The Road to Artificial SuperIntelligence: A Comprehensive Survey of Superalignment
topic Machine Learning
url https://arxiv.org/abs/2412.16468