Improving Neutral Point-of-View Generation with Data- and Parameter-Efficient RL

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hoffmann, Jessica, Ahlheim, Christiane, Yu, Zac, Walfrand, Aria, Jin, Jarvis, Tano, Marie, Beirami, Ahmad, van Liemt, Erin, Thain, Nithum, Sidahmed, Hakim, Dixon, Lucas
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916996700962816
author Hoffmann, Jessica
Ahlheim, Christiane
Yu, Zac
Walfrand, Aria
Jin, Jarvis
Tano, Marie
Beirami, Ahmad
van Liemt, Erin
Thain, Nithum
Sidahmed, Hakim
Dixon, Lucas
author_facet Hoffmann, Jessica
Ahlheim, Christiane
Yu, Zac
Walfrand, Aria
Jin, Jarvis
Tano, Marie
Beirami, Ahmad
van Liemt, Erin
Thain, Nithum
Sidahmed, Hakim
Dixon, Lucas
contents The paper shows that parameter-efficient reinforcement learning (PE-RL) is a highly effective training regime to improve large language models' (LLMs) ability to answer queries on sensitive topics with a Neutral Point of View (NPOV), i.e. to provide significantly more informative, diverse and impartial answers. This is shown by evaluating PE-RL and multiple strong baselines-including LoRA finetuning (strongest baseline), SFT and RLHF. PE-RL not only improves on overall NPOV quality compared to the strongest baseline ($97.06\%\rightarrow 99.08\%$), but also scores much higher on features linguists identify as key to separating sufficient answers from "great'' answers ($60.25\%\rightarrow 85.21\%$ for presence of supportive details, $68.74\%\rightarrow 91.43\%$ for absence of oversimplification). A qualitative analysis corroborates this. Moreover, our evaluation also finds a key property of PE-RL for this task: unlike methods that update all parameters, it generalises out of topic. Finally, to enable further studies we also release the dataset, SHQ-NPOV, and provide a methodology to create such datasets through iterative rounds of human peer-critique and annotator training.
format Preprint
id arxiv_https___arxiv_org_abs_2503_03654
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Improving Neutral Point-of-View Generation with Data- and Parameter-Efficient RL
Hoffmann, Jessica
Ahlheim, Christiane
Yu, Zac
Walfrand, Aria
Jin, Jarvis
Tano, Marie
Beirami, Ahmad
van Liemt, Erin
Thain, Nithum
Sidahmed, Hakim
Dixon, Lucas
Computation and Language
Artificial Intelligence
Machine Learning
The paper shows that parameter-efficient reinforcement learning (PE-RL) is a highly effective training regime to improve large language models' (LLMs) ability to answer queries on sensitive topics with a Neutral Point of View (NPOV), i.e. to provide significantly more informative, diverse and impartial answers. This is shown by evaluating PE-RL and multiple strong baselines-including LoRA finetuning (strongest baseline), SFT and RLHF. PE-RL not only improves on overall NPOV quality compared to the strongest baseline ($97.06\%\rightarrow 99.08\%$), but also scores much higher on features linguists identify as key to separating sufficient answers from "great'' answers ($60.25\%\rightarrow 85.21\%$ for presence of supportive details, $68.74\%\rightarrow 91.43\%$ for absence of oversimplification). A qualitative analysis corroborates this. Moreover, our evaluation also finds a key property of PE-RL for this task: unlike methods that update all parameters, it generalises out of topic. Finally, to enable further studies we also release the dataset, SHQ-NPOV, and provide a methodology to create such datasets through iterative rounds of human peer-critique and annotator training.
title Improving Neutral Point-of-View Generation with Data- and Parameter-Efficient RL
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2503.03654