Uncovering Latent Human Wellbeing in Language Model Embeddings

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Freire, Pedro, Tan, ChengCheng, Gleave, Adam, Hendrycks, Dan, Emmons, Scott
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913237028569088
author Freire, Pedro
Tan, ChengCheng
Gleave, Adam
Hendrycks, Dan
Emmons, Scott
author_facet Freire, Pedro
Tan, ChengCheng
Gleave, Adam
Hendrycks, Dan
Emmons, Scott
contents Do language models implicitly learn a concept of human wellbeing? We explore this through the ETHICS Utilitarianism task, assessing if scaling enhances pretrained models' representations. Our initial finding reveals that, without any prompt engineering or finetuning, the leading principal component from OpenAI's text-embedding-ada-002 achieves 73.9% accuracy. This closely matches the 74.6% of BERT-large finetuned on the entire ETHICS dataset, suggesting pretraining conveys some understanding about human wellbeing. Next, we consider four language model families, observing how Utilitarianism accuracy varies with increased parameters. We find performance is nondecreasing with increased model size when using sufficient numbers of principal components.
format Preprint
id arxiv_https___arxiv_org_abs_2402_11777
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Uncovering Latent Human Wellbeing in Language Model Embeddings
Freire, Pedro
Tan, ChengCheng
Gleave, Adam
Hendrycks, Dan
Emmons, Scott
Computation and Language
Artificial Intelligence
Machine Learning
I.2.7
Do language models implicitly learn a concept of human wellbeing? We explore this through the ETHICS Utilitarianism task, assessing if scaling enhances pretrained models' representations. Our initial finding reveals that, without any prompt engineering or finetuning, the leading principal component from OpenAI's text-embedding-ada-002 achieves 73.9% accuracy. This closely matches the 74.6% of BERT-large finetuned on the entire ETHICS dataset, suggesting pretraining conveys some understanding about human wellbeing. Next, we consider four language model families, observing how Utilitarianism accuracy varies with increased parameters. We find performance is nondecreasing with increased model size when using sufficient numbers of principal components.
title Uncovering Latent Human Wellbeing in Language Model Embeddings
topic Computation and Language
Artificial Intelligence
Machine Learning
I.2.7
url https://arxiv.org/abs/2402.11777