Counterfactual Samples Constructing and Training for Commonsense Statements Estimation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Chong, Feng, Zaiwen, Liu, Lin, Deng, Zhenyun, Li, Jiuyong, Zhai, Ruifang, Cheng, Debo, Qin, Li
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910766388477952
author Liu, Chong
Feng, Zaiwen
Liu, Lin
Deng, Zhenyun
Li, Jiuyong
Zhai, Ruifang
Cheng, Debo
Qin, Li
author_facet Liu, Chong
Feng, Zaiwen
Liu, Lin
Deng, Zhenyun
Li, Jiuyong
Zhai, Ruifang
Cheng, Debo
Qin, Li
contents Plausibility Estimation (PE) plays a crucial role for enabling language models to objectively comprehend the real world. While large language models (LLMs) demonstrate remarkable capabilities in PE tasks but sometimes produce trivial commonsense errors due to the complexity of commonsense knowledge. They lack two key traits of an ideal PE model: a) Language-explainable: relying on critical word segments for decisions, and b) Commonsense-sensitive: detecting subtle linguistic variations in commonsense. To address these issues, we propose a novel model-agnostic method, referred to as Commonsense Counterfactual Samples Generating (CCSG). By training PE models with CCSG, we encourage them to focus on critical words, thereby enhancing both their language-explainable and commonsense-sensitive capabilities. Specifically, CCSG generates counterfactual samples by strategically replacing key words and introducing low-level dropout within sentences. These counterfactual samples are then incorporated into a sentence-level contrastive training framework to further enhance the model's learning process. Experimental results across nine diverse datasets demonstrate the effectiveness of CCSG in addressing commonsense reasoning challenges, with our CCSG method showing 3.07% improvement against the SOTA methods.
format Preprint
id arxiv_https___arxiv_org_abs_2412_20563
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Counterfactual Samples Constructing and Training for Commonsense Statements Estimation
Liu, Chong
Feng, Zaiwen
Liu, Lin
Deng, Zhenyun
Li, Jiuyong
Zhai, Ruifang
Cheng, Debo
Qin, Li
Computation and Language
Plausibility Estimation (PE) plays a crucial role for enabling language models to objectively comprehend the real world. While large language models (LLMs) demonstrate remarkable capabilities in PE tasks but sometimes produce trivial commonsense errors due to the complexity of commonsense knowledge. They lack two key traits of an ideal PE model: a) Language-explainable: relying on critical word segments for decisions, and b) Commonsense-sensitive: detecting subtle linguistic variations in commonsense. To address these issues, we propose a novel model-agnostic method, referred to as Commonsense Counterfactual Samples Generating (CCSG). By training PE models with CCSG, we encourage them to focus on critical words, thereby enhancing both their language-explainable and commonsense-sensitive capabilities. Specifically, CCSG generates counterfactual samples by strategically replacing key words and introducing low-level dropout within sentences. These counterfactual samples are then incorporated into a sentence-level contrastive training framework to further enhance the model's learning process. Experimental results across nine diverse datasets demonstrate the effectiveness of CCSG in addressing commonsense reasoning challenges, with our CCSG method showing 3.07% improvement against the SOTA methods.
title Counterfactual Samples Constructing and Training for Commonsense Statements Estimation
topic Computation and Language
url https://arxiv.org/abs/2412.20563