Self-Consistency Preference Optimization

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Prasad, Archiki, Yuan, Weizhe, Pang, Richard Yuanzhe, Xu, Jing, Fazel-Zarandi, Maryam, Bansal, Mohit, Sukhbaatar, Sainbayar, Weston, Jason, Yu, Jane
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916828013395968
author Prasad, Archiki
Yuan, Weizhe
Pang, Richard Yuanzhe
Xu, Jing
Fazel-Zarandi, Maryam
Bansal, Mohit
Sukhbaatar, Sainbayar
Weston, Jason
Yu, Jane
author_facet Prasad, Archiki
Yuan, Weizhe
Pang, Richard Yuanzhe
Xu, Jing
Fazel-Zarandi, Maryam
Bansal, Mohit
Sukhbaatar, Sainbayar
Weston, Jason
Yu, Jane
contents Self-alignment, whereby models learn to improve themselves without human annotation, is a rapidly growing research area. However, existing techniques often fail to improve complex reasoning tasks due to the difficulty of assigning correct rewards. An orthogonal approach that is known to improve correctness is self-consistency, a method applied at inference time based on multiple sampling in order to find the most consistent answer. In this work, we extend the self-consistency concept to help train models. We thus introduce self-consistency preference optimization (ScPO), which iteratively trains consistent answers to be preferred over inconsistent ones on unsupervised new problems. We show ScPO leads to large improvements over conventional reward model training on reasoning tasks such as GSM8K and MATH, closing the gap with supervised training with gold answers or preferences, and that combining ScPO with standard supervised learning improves results even further. On ZebraLogic, ScPO finetunes Llama-3 8B to be superior to Llama-3 70B, Gemma-2 27B, and Claude-3 Haiku.
format Preprint
id arxiv_https___arxiv_org_abs_2411_04109
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Self-Consistency Preference Optimization
Prasad, Archiki
Yuan, Weizhe
Pang, Richard Yuanzhe
Xu, Jing
Fazel-Zarandi, Maryam
Bansal, Mohit
Sukhbaatar, Sainbayar
Weston, Jason
Yu, Jane
Computation and Language
Artificial Intelligence
Machine Learning
Self-alignment, whereby models learn to improve themselves without human annotation, is a rapidly growing research area. However, existing techniques often fail to improve complex reasoning tasks due to the difficulty of assigning correct rewards. An orthogonal approach that is known to improve correctness is self-consistency, a method applied at inference time based on multiple sampling in order to find the most consistent answer. In this work, we extend the self-consistency concept to help train models. We thus introduce self-consistency preference optimization (ScPO), which iteratively trains consistent answers to be preferred over inconsistent ones on unsupervised new problems. We show ScPO leads to large improvements over conventional reward model training on reasoning tasks such as GSM8K and MATH, closing the gap with supervised training with gold answers or preferences, and that combining ScPO with standard supervised learning improves results even further. On ZebraLogic, ScPO finetunes Llama-3 8B to be superior to Llama-3 70B, Gemma-2 27B, and Claude-3 Haiku.
title Self-Consistency Preference Optimization
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2411.04109