Rationales Are Not Silver Bullets: Measuring the Impact of Rationales on Model Performance and Reliability

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Chiwei, Xu, Benfeng, Yang, An, Lin, Junyang, Wang, Quan, Zhou, Chang, Mao, Zhendong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913866778148864
author Zhu, Chiwei
Xu, Benfeng
Yang, An
Lin, Junyang
Wang, Quan
Zhou, Chang
Mao, Zhendong
author_facet Zhu, Chiwei
Xu, Benfeng
Yang, An
Lin, Junyang
Wang, Quan
Zhou, Chang
Mao, Zhendong
contents Training language models with rationales augmentation has been shown to be beneficial in many existing works. In this paper, we identify that such a prevailing view does not hold consistently. We conduct comprehensive investigations to thoroughly inspect the impact of rationales on model performance as well as a novel perspective of model reliability. The results lead to several key findings that add new insights upon existing understandings: 1) Rationales can, at times, deteriorate model performance; 2) Rationales can, at times, improve model reliability, even outperforming their untrained counterparts; 3) A linear correspondence exists in between the performance and reliability improvements, while both are driven by the intrinsic difficulty of the task. These findings provide informative regulations on the broad utilization of rationales and raise critical implications on the procedure of explicitly aligning language models with implicit human thoughts. Codes can be found at https://github.com/Ignoramus0817/rationales.
format Preprint
id arxiv_https___arxiv_org_abs_2505_24147
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Rationales Are Not Silver Bullets: Measuring the Impact of Rationales on Model Performance and Reliability
Zhu, Chiwei
Xu, Benfeng
Yang, An
Lin, Junyang
Wang, Quan
Zhou, Chang
Mao, Zhendong
Computation and Language
Training language models with rationales augmentation has been shown to be beneficial in many existing works. In this paper, we identify that such a prevailing view does not hold consistently. We conduct comprehensive investigations to thoroughly inspect the impact of rationales on model performance as well as a novel perspective of model reliability. The results lead to several key findings that add new insights upon existing understandings: 1) Rationales can, at times, deteriorate model performance; 2) Rationales can, at times, improve model reliability, even outperforming their untrained counterparts; 3) A linear correspondence exists in between the performance and reliability improvements, while both are driven by the intrinsic difficulty of the task. These findings provide informative regulations on the broad utilization of rationales and raise critical implications on the procedure of explicitly aligning language models with implicit human thoughts. Codes can be found at https://github.com/Ignoramus0817/rationales.
title Rationales Are Not Silver Bullets: Measuring the Impact of Rationales on Model Performance and Reliability
topic Computation and Language
url https://arxiv.org/abs/2505.24147