Self-rationalization improves LLM as a fine-grained judge

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Trivedi, Prapti, Gulati, Aditya, Molenschot, Oliver, Rajeev, Meghana Arakkal, Ramamurthy, Rajkumar, Stevens, Keith, Chaudhery, Tanveesh Singh, Jambholkar, Jahnavi, Zou, James, Rajani, Nazneen
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929531286192128
author Trivedi, Prapti
Gulati, Aditya
Molenschot, Oliver
Rajeev, Meghana Arakkal
Ramamurthy, Rajkumar
Stevens, Keith
Chaudhery, Tanveesh Singh
Jambholkar, Jahnavi
Zou, James
Rajani, Nazneen
author_facet Trivedi, Prapti
Gulati, Aditya
Molenschot, Oliver
Rajeev, Meghana Arakkal
Ramamurthy, Rajkumar
Stevens, Keith
Chaudhery, Tanveesh Singh
Jambholkar, Jahnavi
Zou, James
Rajani, Nazneen
contents LLM-as-a-judge models have been used for evaluating both human and AI generated content, specifically by providing scores and rationales. Rationales, in addition to increasing transparency, help models learn to calibrate its judgments. Enhancing a model's rationale can therefore improve its calibration abilities and ultimately the ability to score content. We introduce Self-Rationalization, an iterative process of improving the rationales for the judge models, which consequently improves the score for fine-grained customizable scoring criteria (i.e., likert-scale scoring with arbitrary evaluation criteria). Self-rationalization works by having the model generate multiple judgments with rationales for the same input, curating a preference pair dataset from its own judgements, and iteratively fine-tuning the judge via DPO. Intuitively, this approach allows the judge model to self-improve by learning from its own rationales, leading to better alignment and evaluation accuracy. After just two iterations -- while only relying on examples in the training set -- human evaluation shows that our judge model learns to produce higher quality rationales, with a win rate of $62\%$ on average compared to models just trained via SFT on rationale . This judge model also achieves high scoring accuracy on BigGen Bench and Reward Bench, outperforming even bigger sized models trained using SFT with rationale, self-consistency or best-of-$N$ sampling by $3\%$ to $9\%$.
format Preprint
id arxiv_https___arxiv_org_abs_2410_05495
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Self-rationalization improves LLM as a fine-grained judge
Trivedi, Prapti
Gulati, Aditya
Molenschot, Oliver
Rajeev, Meghana Arakkal
Ramamurthy, Rajkumar
Stevens, Keith
Chaudhery, Tanveesh Singh
Jambholkar, Jahnavi
Zou, James
Rajani, Nazneen
Computation and Language
LLM-as-a-judge models have been used for evaluating both human and AI generated content, specifically by providing scores and rationales. Rationales, in addition to increasing transparency, help models learn to calibrate its judgments. Enhancing a model's rationale can therefore improve its calibration abilities and ultimately the ability to score content. We introduce Self-Rationalization, an iterative process of improving the rationales for the judge models, which consequently improves the score for fine-grained customizable scoring criteria (i.e., likert-scale scoring with arbitrary evaluation criteria). Self-rationalization works by having the model generate multiple judgments with rationales for the same input, curating a preference pair dataset from its own judgements, and iteratively fine-tuning the judge via DPO. Intuitively, this approach allows the judge model to self-improve by learning from its own rationales, leading to better alignment and evaluation accuracy. After just two iterations -- while only relying on examples in the training set -- human evaluation shows that our judge model learns to produce higher quality rationales, with a win rate of $62\%$ on average compared to models just trained via SFT on rationale . This judge model also achieves high scoring accuracy on BigGen Bench and Reward Bench, outperforming even bigger sized models trained using SFT with rationale, self-consistency or best-of-$N$ sampling by $3\%$ to $9\%$.
title Self-rationalization improves LLM as a fine-grained judge
topic Computation and Language
url https://arxiv.org/abs/2410.05495