Compartmentalised Agentic Reasoning for Clinical NLI

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jullien, Maël, Xu, Lei, Valentino, Marco, Freitas, André
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908768272384000
author Jullien, Maël
Xu, Lei
Valentino, Marco
Freitas, André
author_facet Jullien, Maël
Xu, Lei
Valentino, Marco
Freitas, André
contents Large language models can produce fluent judgments for clinical natural language inference, yet they frequently fail when the decision requires the correct inferential schema rather than surface matching. We introduce CARENLI, a compartmentalised agentic framework that routes each premise-statement pair to a reasoning family and then applies a specialised solver with explicit verification and targeted refinement. We evaluate on an expanded CTNLI benchmark of 200 instances spanning four reasoning families: Causal Attribution, Compositional Grounding, Epistemic Verification, and Risk State Abstraction. Across four contemporary backbone models, CARENLI improves mean accuracy from about 23% with direct prompting to about 57%, a gain of roughly 34 points, with the largest benefits on structurally demanding reasoning types. These results support compartmentalisation plus verification as a practical route to more reliable and auditable clinical inference.
format Preprint
id arxiv_https___arxiv_org_abs_2509_10222
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Compartmentalised Agentic Reasoning for Clinical NLI
Jullien, Maël
Xu, Lei
Valentino, Marco
Freitas, André
Artificial Intelligence
Large language models can produce fluent judgments for clinical natural language inference, yet they frequently fail when the decision requires the correct inferential schema rather than surface matching. We introduce CARENLI, a compartmentalised agentic framework that routes each premise-statement pair to a reasoning family and then applies a specialised solver with explicit verification and targeted refinement. We evaluate on an expanded CTNLI benchmark of 200 instances spanning four reasoning families: Causal Attribution, Compositional Grounding, Epistemic Verification, and Risk State Abstraction. Across four contemporary backbone models, CARENLI improves mean accuracy from about 23% with direct prompting to about 57%, a gain of roughly 34 points, with the largest benefits on structurally demanding reasoning types. These results support compartmentalisation plus verification as a practical route to more reliable and auditable clinical inference.
title Compartmentalised Agentic Reasoning for Clinical NLI
topic Artificial Intelligence
url https://arxiv.org/abs/2509.10222