Improving and Assessing the Fidelity of Large Language Models Alignment to Online Communities

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chu, Minh Duc, He, Zihao, Dorn, Rebecca, Lerman, Kristina
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913685646082048
author Chu, Minh Duc
He, Zihao
Dorn, Rebecca
Lerman, Kristina
author_facet Chu, Minh Duc
He, Zihao
Dorn, Rebecca
Lerman, Kristina
contents Large language models (LLMs) have shown promise in representing individuals and communities, offering new ways to study complex social dynamics. However, effectively aligning LLMs with specific human groups and systematically assessing the fidelity of the alignment remains a challenge. This paper presents a robust framework for aligning LLMs with online communities via instruction-tuning and comprehensively evaluating alignment across various aspects of language, including authenticity, emotional tone, toxicity, and harm. We demonstrate the utility of our approach by applying it to online communities centered on dieting and body image. We administer an eating disorder psychometric test to the aligned LLMs to reveal unhealthy beliefs and successfully differentiate communities with varying levels of eating disorder risk. Our results highlight the potential of LLMs in automated moderation and broader applications in public health and social science research.
format Preprint
id arxiv_https___arxiv_org_abs_2408_09366
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Improving and Assessing the Fidelity of Large Language Models Alignment to Online Communities
Chu, Minh Duc
He, Zihao
Dorn, Rebecca
Lerman, Kristina
Computation and Language
Computers and Society
Social and Information Networks
Large language models (LLMs) have shown promise in representing individuals and communities, offering new ways to study complex social dynamics. However, effectively aligning LLMs with specific human groups and systematically assessing the fidelity of the alignment remains a challenge. This paper presents a robust framework for aligning LLMs with online communities via instruction-tuning and comprehensively evaluating alignment across various aspects of language, including authenticity, emotional tone, toxicity, and harm. We demonstrate the utility of our approach by applying it to online communities centered on dieting and body image. We administer an eating disorder psychometric test to the aligned LLMs to reveal unhealthy beliefs and successfully differentiate communities with varying levels of eating disorder risk. Our results highlight the potential of LLMs in automated moderation and broader applications in public health and social science research.
title Improving and Assessing the Fidelity of Large Language Models Alignment to Online Communities
topic Computation and Language
Computers and Society
Social and Information Networks
url https://arxiv.org/abs/2408.09366