Calibrating LLMs with Preference Optimization on Thought Trees for Generating Rationale in Science Question Scoring

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Li, Jiazheng, Xu, Hainiu, Sun, Zhaoyue, Zhou, Yuxiang, West, David, Aloisi, Cesare, He, Yulan
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913542584664064
author Li, Jiazheng
Xu, Hainiu
Sun, Zhaoyue
Zhou, Yuxiang
West, David
Aloisi, Cesare
He, Yulan
author_facet Li, Jiazheng
Xu, Hainiu
Sun, Zhaoyue
Zhou, Yuxiang
West, David
Aloisi, Cesare
He, Yulan
contents Generating rationales that justify scoring decisions has been a promising way to facilitate explainability in automated scoring systems. However, existing methods do not match the accuracy of classifier-based methods. Plus, the generated rationales often contain hallucinated information. To address these issues, we propose a novel framework capable of generating more faithful rationales and, more importantly, matching performance with classifier-based black-box scoring systems. We first mimic the human assessment process by querying Large Language Models (LLMs) to generate a thought tree. We then summarise intermediate assessment decisions from each thought tree path for creating synthetic rationale data and rationale preference data. Finally, we utilise the generated synthetic data to calibrate LLMs through a two-step training process: supervised fine-tuning and preference optimization. Extensive experimental results demonstrate that our framework achieves a 38% assessment performance improvement in the QWK score compared to prior work while producing higher-quality rationales, as recognised by human evaluators and LLMs. Our work sheds light on the effectiveness of performing preference optimization using synthetic preference data obtained from thought tree paths. Data and code are available at https://github.com/lijiazheng99/thought_tree_assessment.
format Preprint
id arxiv_https___arxiv_org_abs_2406_19949
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Calibrating LLMs with Preference Optimization on Thought Trees for Generating Rationale in Science Question Scoring
Li, Jiazheng
Xu, Hainiu
Sun, Zhaoyue
Zhou, Yuxiang
West, David
Aloisi, Cesare
He, Yulan
Computation and Language
Generating rationales that justify scoring decisions has been a promising way to facilitate explainability in automated scoring systems. However, existing methods do not match the accuracy of classifier-based methods. Plus, the generated rationales often contain hallucinated information. To address these issues, we propose a novel framework capable of generating more faithful rationales and, more importantly, matching performance with classifier-based black-box scoring systems. We first mimic the human assessment process by querying Large Language Models (LLMs) to generate a thought tree. We then summarise intermediate assessment decisions from each thought tree path for creating synthetic rationale data and rationale preference data. Finally, we utilise the generated synthetic data to calibrate LLMs through a two-step training process: supervised fine-tuning and preference optimization. Extensive experimental results demonstrate that our framework achieves a 38% assessment performance improvement in the QWK score compared to prior work while producing higher-quality rationales, as recognised by human evaluators and LLMs. Our work sheds light on the effectiveness of performing preference optimization using synthetic preference data obtained from thought tree paths. Data and code are available at https://github.com/lijiazheng99/thought_tree_assessment.
title Calibrating LLMs with Preference Optimization on Thought Trees for Generating Rationale in Science Question Scoring
topic Computation and Language
url https://arxiv.org/abs/2406.19949