Towards Scalable Automated Alignment of LLMs: A Survey

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cao, Boxi, Lu, Keming, Lu, Xinyu, Chen, Jiawei, Ren, Mengjie, Xiang, Hao, Liu, Peilin, Lu, Yaojie, He, Ben, Han, Xianpei, Sun, Le, Lin, Hongyu, Yu, Bowen
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917767000621056
author Cao, Boxi
Lu, Keming
Lu, Xinyu
Chen, Jiawei
Ren, Mengjie
Xiang, Hao
Liu, Peilin
Lu, Yaojie
He, Ben
Han, Xianpei
Sun, Le
Lin, Hongyu
Yu, Bowen
author_facet Cao, Boxi
Lu, Keming
Lu, Xinyu
Chen, Jiawei
Ren, Mengjie
Xiang, Hao
Liu, Peilin
Lu, Yaojie
He, Ben
Han, Xianpei
Sun, Le
Lin, Hongyu
Yu, Bowen
contents Alignment is the most critical step in building large language models (LLMs) that meet human needs. With the rapid development of LLMs gradually surpassing human capabilities, traditional alignment methods based on human-annotation are increasingly unable to meet the scalability demands. Therefore, there is an urgent need to explore new sources of automated alignment signals and technical approaches. In this paper, we systematically review the recently emerging methods of automated alignment, attempting to explore how to achieve effective, scalable, automated alignment once the capabilities of LLMs exceed those of humans. Specifically, we categorize existing automated alignment methods into 4 major categories based on the sources of alignment signals and discuss the current status and potential development of each category. Additionally, we explore the underlying mechanisms that enable automated alignment and discuss the essential factors that make automated alignment technologies feasible and effective from the fundamental role of alignment.
format Preprint
id arxiv_https___arxiv_org_abs_2406_01252
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Towards Scalable Automated Alignment of LLMs: A Survey
Cao, Boxi
Lu, Keming
Lu, Xinyu
Chen, Jiawei
Ren, Mengjie
Xiang, Hao
Liu, Peilin
Lu, Yaojie
He, Ben
Han, Xianpei
Sun, Le
Lin, Hongyu
Yu, Bowen
Computation and Language
Artificial Intelligence
Machine Learning
Alignment is the most critical step in building large language models (LLMs) that meet human needs. With the rapid development of LLMs gradually surpassing human capabilities, traditional alignment methods based on human-annotation are increasingly unable to meet the scalability demands. Therefore, there is an urgent need to explore new sources of automated alignment signals and technical approaches. In this paper, we systematically review the recently emerging methods of automated alignment, attempting to explore how to achieve effective, scalable, automated alignment once the capabilities of LLMs exceed those of humans. Specifically, we categorize existing automated alignment methods into 4 major categories based on the sources of alignment signals and discuss the current status and potential development of each category. Additionally, we explore the underlying mechanisms that enable automated alignment and discuss the essential factors that make automated alignment technologies feasible and effective from the fundamental role of alignment.
title Towards Scalable Automated Alignment of LLMs: A Survey
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2406.01252