Reasoning Elicitation in Language Models via Counterfactual Feedback

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hüyük, Alihan, Xu, Xinnuo, Maasch, Jacqueline, Nori, Aditya V., González, Javier
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929760847790080
author Hüyük, Alihan
Xu, Xinnuo
Maasch, Jacqueline
Nori, Aditya V.
González, Javier
author_facet Hüyük, Alihan
Xu, Xinnuo
Maasch, Jacqueline
Nori, Aditya V.
González, Javier
contents Despite the increasing effectiveness of language models, their reasoning capabilities remain underdeveloped. In particular, causal reasoning through counterfactual question answering is lacking. This work aims to bridge this gap. We first derive novel metrics that balance accuracy in factual and counterfactual questions, capturing a more complete view of the reasoning abilities of language models than traditional factual-only based metrics. Second, we propose several fine-tuning approaches that aim to elicit better reasoning mechanisms, in the sense of the proposed metrics. Finally, we evaluate the performance of the fine-tuned language models in a variety of realistic scenarios. In particular, we investigate to what extent our fine-tuning approaches systemically achieve better generalization with respect to the base models in several problems that require, among others, inductive and deductive reasoning capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2410_03767
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Reasoning Elicitation in Language Models via Counterfactual Feedback
Hüyük, Alihan
Xu, Xinnuo
Maasch, Jacqueline
Nori, Aditya V.
González, Javier
Computation and Language
Artificial Intelligence
Machine Learning
Despite the increasing effectiveness of language models, their reasoning capabilities remain underdeveloped. In particular, causal reasoning through counterfactual question answering is lacking. This work aims to bridge this gap. We first derive novel metrics that balance accuracy in factual and counterfactual questions, capturing a more complete view of the reasoning abilities of language models than traditional factual-only based metrics. Second, we propose several fine-tuning approaches that aim to elicit better reasoning mechanisms, in the sense of the proposed metrics. Finally, we evaluate the performance of the fine-tuned language models in a variety of realistic scenarios. In particular, we investigate to what extent our fine-tuning approaches systemically achieve better generalization with respect to the base models in several problems that require, among others, inductive and deductive reasoning capabilities.
title Reasoning Elicitation in Language Models via Counterfactual Feedback
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2410.03767