Measuring What LLMs Think They Do: SHAP Faithfulness and Deployability on Financial Tabular Classification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: AlMarri, Saeed, Ravaut, Mathieu, Juhasz, Kristof, Marti, Gautier, Ahbabi, Hamdan Al, Elfadel, Ibrahim
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914174493261824
author AlMarri, Saeed
Ravaut, Mathieu
Juhasz, Kristof
Marti, Gautier
Ahbabi, Hamdan Al
Elfadel, Ibrahim
author_facet AlMarri, Saeed
Ravaut, Mathieu
Juhasz, Kristof
Marti, Gautier
Ahbabi, Hamdan Al
Elfadel, Ibrahim
contents Large Language Models (LLMs) have attracted significant attention for classification tasks, offering a flexible alternative to trusted classical machine learning models like LightGBM through zero-shot prompting. However, their reliability for structured tabular data remains unclear, particularly in high stakes applications like financial risk assessment. Our study systematically evaluates LLMs and generates their SHAP values on financial classification tasks. Our analysis shows a divergence between LLMs self-explanation of feature impact and their SHAP values, as well as notable differences between LLMs and LightGBM SHAP values. These findings highlight the limitations of LLMs as standalone classifiers for structured financial modeling, but also instill optimism that improved explainability mechanisms coupled with few-shot prompting will make LLMs usable in risk-sensitive domains.
format Preprint
id arxiv_https___arxiv_org_abs_2512_00163
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Measuring What LLMs Think They Do: SHAP Faithfulness and Deployability on Financial Tabular Classification
AlMarri, Saeed
Ravaut, Mathieu
Juhasz, Kristof
Marti, Gautier
Ahbabi, Hamdan Al
Elfadel, Ibrahim
Machine Learning
Computation and Language
Large Language Models (LLMs) have attracted significant attention for classification tasks, offering a flexible alternative to trusted classical machine learning models like LightGBM through zero-shot prompting. However, their reliability for structured tabular data remains unclear, particularly in high stakes applications like financial risk assessment. Our study systematically evaluates LLMs and generates their SHAP values on financial classification tasks. Our analysis shows a divergence between LLMs self-explanation of feature impact and their SHAP values, as well as notable differences between LLMs and LightGBM SHAP values. These findings highlight the limitations of LLMs as standalone classifiers for structured financial modeling, but also instill optimism that improved explainability mechanisms coupled with few-shot prompting will make LLMs usable in risk-sensitive domains.
title Measuring What LLMs Think They Do: SHAP Faithfulness and Deployability on Financial Tabular Classification
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2512.00163