Filling the Gap: Is Commonsense Knowledge Generation useful for Natural Language Inference?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jayaweera, Chathuri, Yanqui, Brianna, Dorr, Bonnie
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908784272605184
author Jayaweera, Chathuri
Yanqui, Brianna
Dorr, Bonnie
author_facet Jayaweera, Chathuri
Yanqui, Brianna
Dorr, Bonnie
contents Natural Language Inference (NLI) is the task of determining whether a premise entails, contradicts, or is neutral with respect to a given hypothesis. The task is often framed as emulating human inferential processes, in which commonsense knowledge plays a major role. This study examines whether Large Language Models (LLMs) can generate useful commonsense axioms for Natural Language Inference, and evaluates their impact on performance using the SNLI and ANLI benchmarks with the Llama-3.1-70B and gpt-oss-120b models. We show that a hybrid approach, which selectively provides highly factual axioms based on judged helpfulness, yields consistent accuracy improvements of 1.99% to 6.88% across tested configurations, demonstrating the effectiveness of selective knowledge access for NLI. We also find that this targeted use of commonsense knowledge helps models overcome a bias toward the Neutral class by providing essential real-world context.
format Preprint
id arxiv_https___arxiv_org_abs_2507_15100
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Filling the Gap: Is Commonsense Knowledge Generation useful for Natural Language Inference?
Jayaweera, Chathuri
Yanqui, Brianna
Dorr, Bonnie
Computation and Language
Artificial Intelligence
Natural Language Inference (NLI) is the task of determining whether a premise entails, contradicts, or is neutral with respect to a given hypothesis. The task is often framed as emulating human inferential processes, in which commonsense knowledge plays a major role. This study examines whether Large Language Models (LLMs) can generate useful commonsense axioms for Natural Language Inference, and evaluates their impact on performance using the SNLI and ANLI benchmarks with the Llama-3.1-70B and gpt-oss-120b models. We show that a hybrid approach, which selectively provides highly factual axioms based on judged helpfulness, yields consistent accuracy improvements of 1.99% to 6.88% across tested configurations, demonstrating the effectiveness of selective knowledge access for NLI. We also find that this targeted use of commonsense knowledge helps models overcome a bias toward the Neutral class by providing essential real-world context.
title Filling the Gap: Is Commonsense Knowledge Generation useful for Natural Language Inference?
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2507.15100