Towards Consistent Natural-Language Explanations via Explanation-Consistency Finetuning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Yanda, Singh, Chandan, Liu, Xiaodong, Zuo, Simiao, Yu, Bin, He, He, Gao, Jianfeng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914652746678272
author Chen, Yanda
Singh, Chandan
Liu, Xiaodong
Zuo, Simiao
Yu, Bin
He, He
Gao, Jianfeng
author_facet Chen, Yanda
Singh, Chandan
Liu, Xiaodong
Zuo, Simiao
Yu, Bin
He, He
Gao, Jianfeng
contents Large language models (LLMs) often generate convincing, fluent explanations. However, different from humans, they often generate inconsistent explanations on different inputs. For example, an LLM may generate the explanation "all birds can fly" when answering the question "Can sparrows fly?" but meanwhile answer "no" to the related question "Can penguins fly?". Explanations should be consistent across related examples so that they allow a human to simulate the LLM's decision process on multiple examples. We propose explanation-consistency finetuning (EC-finetuning), a method that adapts LLMs to generate more consistent natural-language explanations on related examples. EC-finetuning involves finetuning LLMs on synthetic data that is carefully constructed to contain consistent explanations. Across a variety of question-answering datasets in various domains, EC-finetuning yields a 10.0% relative explanation consistency improvement on four finetuning datasets, and generalizes to seven out-of-distribution datasets not seen during finetuning (+4.5% relative). Code is available at https://github.com/yandachen/explanation-consistency-finetuning .
format Preprint
id arxiv_https___arxiv_org_abs_2401_13986
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Towards Consistent Natural-Language Explanations via Explanation-Consistency Finetuning
Chen, Yanda
Singh, Chandan
Liu, Xiaodong
Zuo, Simiao
Yu, Bin
He, He
Gao, Jianfeng
Computation and Language
Artificial Intelligence
Machine Learning
Large language models (LLMs) often generate convincing, fluent explanations. However, different from humans, they often generate inconsistent explanations on different inputs. For example, an LLM may generate the explanation "all birds can fly" when answering the question "Can sparrows fly?" but meanwhile answer "no" to the related question "Can penguins fly?". Explanations should be consistent across related examples so that they allow a human to simulate the LLM's decision process on multiple examples. We propose explanation-consistency finetuning (EC-finetuning), a method that adapts LLMs to generate more consistent natural-language explanations on related examples. EC-finetuning involves finetuning LLMs on synthetic data that is carefully constructed to contain consistent explanations. Across a variety of question-answering datasets in various domains, EC-finetuning yields a 10.0% relative explanation consistency improvement on four finetuning datasets, and generalizes to seven out-of-distribution datasets not seen during finetuning (+4.5% relative). Code is available at https://github.com/yandachen/explanation-consistency-finetuning .
title Towards Consistent Natural-Language Explanations via Explanation-Consistency Finetuning
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2401.13986