Better Alignment with Instruction Back-and-Forth Translation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nguyen, Thao, Li, Jeffrey, Oh, Sewoong, Schmidt, Ludwig, Weston, Jason, Zettlemoyer, Luke, Li, Xian
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914911721881600
author Nguyen, Thao
Li, Jeffrey
Oh, Sewoong
Schmidt, Ludwig
Weston, Jason
Zettlemoyer, Luke
Li, Xian
author_facet Nguyen, Thao
Li, Jeffrey
Oh, Sewoong
Schmidt, Ludwig
Weston, Jason
Zettlemoyer, Luke
Li, Xian
contents We propose a new method, instruction back-and-forth translation, to construct high-quality synthetic data grounded in world knowledge for aligning large language models (LLMs). Given documents from a web corpus, we generate and curate synthetic instructions using the backtranslation approach proposed by Li et al.(2023a), and rewrite the responses to improve their quality further based on the initial documents. Fine-tuning with the resulting (backtranslated instruction, rewritten response) pairs yields higher win rates on AlpacaEval than using other common instruction datasets such as Humpback, ShareGPT, Open Orca, Alpaca-GPT4 and Self-instruct. We also demonstrate that rewriting the responses with an LLM outperforms direct distillation, and the two generated text distributions exhibit significant distinction in embedding space. Further analysis shows that our backtranslated instructions are of higher quality than other sources of synthetic instructions, while our responses are more diverse and complex than those obtained from distillation. Overall we find that instruction back-and-forth translation combines the best of both worlds -- making use of the information diversity and quantity found on the web, while ensuring the quality of the responses which is necessary for effective alignment.
format Preprint
id arxiv_https___arxiv_org_abs_2408_04614
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Better Alignment with Instruction Back-and-Forth Translation
Nguyen, Thao
Li, Jeffrey
Oh, Sewoong
Schmidt, Ludwig
Weston, Jason
Zettlemoyer, Luke
Li, Xian
Computation and Language
Artificial Intelligence
Machine Learning
We propose a new method, instruction back-and-forth translation, to construct high-quality synthetic data grounded in world knowledge for aligning large language models (LLMs). Given documents from a web corpus, we generate and curate synthetic instructions using the backtranslation approach proposed by Li et al.(2023a), and rewrite the responses to improve their quality further based on the initial documents. Fine-tuning with the resulting (backtranslated instruction, rewritten response) pairs yields higher win rates on AlpacaEval than using other common instruction datasets such as Humpback, ShareGPT, Open Orca, Alpaca-GPT4 and Self-instruct. We also demonstrate that rewriting the responses with an LLM outperforms direct distillation, and the two generated text distributions exhibit significant distinction in embedding space. Further analysis shows that our backtranslated instructions are of higher quality than other sources of synthetic instructions, while our responses are more diverse and complex than those obtained from distillation. Overall we find that instruction back-and-forth translation combines the best of both worlds -- making use of the information diversity and quantity found on the web, while ensuring the quality of the responses which is necessary for effective alignment.
title Better Alignment with Instruction Back-and-Forth Translation
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2408.04614