Evaluating the role of `Constitutions' for learning from AI feedback

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Redgate, Saskia, Bean, Andrew M., Mahdi, Adam
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929592571265024
author Redgate, Saskia
Bean, Andrew M.
Mahdi, Adam
author_facet Redgate, Saskia
Bean, Andrew M.
Mahdi, Adam
contents The growing capabilities of large language models (LLMs) have led to their use as substitutes for human feedback for training and assessing other LLMs. These methods often rely on `constitutions', written guidelines which a critic model uses to provide feedback and improve generations. We investigate how the choice of constitution affects feedback quality by using four different constitutions to improve patient-centered communication in medical interviews. In pairwise comparisons conducted by 215 human raters, we found that detailed constitutions led to better results regarding emotive qualities. However, none of the constitutions outperformed the baseline in learning more practically-oriented skills related to information gathering and provision. Our findings indicate that while detailed constitutions should be prioritised, there are possible limitations to the effectiveness of AI feedback as a reward signal in certain areas.
format Preprint
id arxiv_https___arxiv_org_abs_2411_10168
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Evaluating the role of `Constitutions' for learning from AI feedback
Redgate, Saskia
Bean, Andrew M.
Mahdi, Adam
Artificial Intelligence
Computation and Language
The growing capabilities of large language models (LLMs) have led to their use as substitutes for human feedback for training and assessing other LLMs. These methods often rely on `constitutions', written guidelines which a critic model uses to provide feedback and improve generations. We investigate how the choice of constitution affects feedback quality by using four different constitutions to improve patient-centered communication in medical interviews. In pairwise comparisons conducted by 215 human raters, we found that detailed constitutions led to better results regarding emotive qualities. However, none of the constitutions outperformed the baseline in learning more practically-oriented skills related to information gathering and provision. Our findings indicate that while detailed constitutions should be prioritised, there are possible limitations to the effectiveness of AI feedback as a reward signal in certain areas.
title Evaluating the role of `Constitutions' for learning from AI feedback
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2411.10168