Unsupervised Elicitation of Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wen, Jiaxin, Ankner, Zachary, Somani, Arushi, Hase, Peter, Marks, Samuel, Goldman-Wetzler, Jacob, Petrini, Linda, Sleight, Henry, Burns, Collin, He, He, Feng, Shi, Perez, Ethan, Leike, Jan
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908789232369664
author Wen, Jiaxin
Ankner, Zachary
Somani, Arushi
Hase, Peter
Marks, Samuel
Goldman-Wetzler, Jacob
Petrini, Linda
Sleight, Henry
Burns, Collin
He, He
Feng, Shi
Perez, Ethan
Leike, Jan
author_facet Wen, Jiaxin
Ankner, Zachary
Somani, Arushi
Hase, Peter
Marks, Samuel
Goldman-Wetzler, Jacob
Petrini, Linda
Sleight, Henry
Burns, Collin
He, He
Feng, Shi
Perez, Ethan
Leike, Jan
contents To steer pretrained language models for downstream tasks, today's post-training paradigm relies on humans to specify desired behaviors. However, for models with superhuman capabilities, it is difficult or impossible to get high-quality human supervision. To address this challenge, we introduce a new unsupervised algorithm, Internal Coherence Maximization (ICM), to fine-tune pretrained language models on their own generated labels, \emph{without external supervision}. On GSM8k-verification, TruthfulQA, and Alpaca reward modeling tasks, our method matches the performance of training on golden labels and outperforms training on crowdsourced human supervision. On tasks where LMs' capabilities are strongly superhuman, our method can elicit those capabilities significantly better than training on human labels. Finally, we show that our method can improve the training of frontier LMs: we use our method to train an unsupervised reward model and use reinforcement learning to train a Claude 4 Sonnet-based assistant. The resulting assistant matches its counterpart trained on production-grade human labels on average, with higher scores on chat and safety yet lower scores on math and coding.
format Preprint
id arxiv_https___arxiv_org_abs_2506_10139
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Unsupervised Elicitation of Language Models
Wen, Jiaxin
Ankner, Zachary
Somani, Arushi
Hase, Peter
Marks, Samuel
Goldman-Wetzler, Jacob
Petrini, Linda
Sleight, Henry
Burns, Collin
He, He
Feng, Shi
Perez, Ethan
Leike, Jan
Computation and Language
Artificial Intelligence
To steer pretrained language models for downstream tasks, today's post-training paradigm relies on humans to specify desired behaviors. However, for models with superhuman capabilities, it is difficult or impossible to get high-quality human supervision. To address this challenge, we introduce a new unsupervised algorithm, Internal Coherence Maximization (ICM), to fine-tune pretrained language models on their own generated labels, \emph{without external supervision}. On GSM8k-verification, TruthfulQA, and Alpaca reward modeling tasks, our method matches the performance of training on golden labels and outperforms training on crowdsourced human supervision. On tasks where LMs' capabilities are strongly superhuman, our method can elicit those capabilities significantly better than training on human labels. Finally, we show that our method can improve the training of frontier LMs: we use our method to train an unsupervised reward model and use reinforcement learning to train a Claude 4 Sonnet-based assistant. The resulting assistant matches its counterpart trained on production-grade human labels on average, with higher scores on chat and safety yet lower scores on math and coding.
title Unsupervised Elicitation of Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2506.10139