A MISMATCHED Benchmark for Scientific Natural Language Inference

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shaik, Firoz, Sadat, Mobashir, Gautam, Nikita, Caragea, Doina, Caragea, Cornelia
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912414836981760
author Shaik, Firoz
Sadat, Mobashir
Gautam, Nikita
Caragea, Doina
Caragea, Cornelia
author_facet Shaik, Firoz
Sadat, Mobashir
Gautam, Nikita
Caragea, Doina
Caragea, Cornelia
contents Scientific Natural Language Inference (NLI) is the task of predicting the semantic relation between a pair of sentences extracted from research articles. Existing datasets for this task are derived from various computer science (CS) domains, whereas non-CS domains are completely ignored. In this paper, we introduce a novel evaluation benchmark for scientific NLI, called MISMATCHED. The new MISMATCHED benchmark covers three non-CS domains-PSYCHOLOGY, ENGINEERING, and PUBLIC HEALTH, and contains 2,700 human annotated sentence pairs. We establish strong baselines on MISMATCHED using both Pre-trained Small Language Models (SLMs) and Large Language Models (LLMs). Our best performing baseline shows a Macro F1 of only 78.17% illustrating the substantial headroom for future improvements. In addition to introducing the MISMATCHED benchmark, we show that incorporating sentence pairs having an implicit scientific NLI relation between them in model training improves their performance on scientific NLI. We make our dataset and code publicly available on GitHub.
format Preprint
id arxiv_https___arxiv_org_abs_2506_04603
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A MISMATCHED Benchmark for Scientific Natural Language Inference
Shaik, Firoz
Sadat, Mobashir
Gautam, Nikita
Caragea, Doina
Caragea, Cornelia
Computation and Language
Scientific Natural Language Inference (NLI) is the task of predicting the semantic relation between a pair of sentences extracted from research articles. Existing datasets for this task are derived from various computer science (CS) domains, whereas non-CS domains are completely ignored. In this paper, we introduce a novel evaluation benchmark for scientific NLI, called MISMATCHED. The new MISMATCHED benchmark covers three non-CS domains-PSYCHOLOGY, ENGINEERING, and PUBLIC HEALTH, and contains 2,700 human annotated sentence pairs. We establish strong baselines on MISMATCHED using both Pre-trained Small Language Models (SLMs) and Large Language Models (LLMs). Our best performing baseline shows a Macro F1 of only 78.17% illustrating the substantial headroom for future improvements. In addition to introducing the MISMATCHED benchmark, we show that incorporating sentence pairs having an implicit scientific NLI relation between them in model training improves their performance on scientific NLI. We make our dataset and code publicly available on GitHub.
title A MISMATCHED Benchmark for Scientific Natural Language Inference
topic Computation and Language
url https://arxiv.org/abs/2506.04603