The Effect of Model Size on LLM Post-hoc Explainability via LIME

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Heyen, Henning, Widdicombe, Amy, Siegel, Noah Y., Perez-Ortiz, Maria, Treleaven, Philip
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916239448735744
author Heyen, Henning
Widdicombe, Amy
Siegel, Noah Y.
Perez-Ortiz, Maria
Treleaven, Philip
author_facet Heyen, Henning
Widdicombe, Amy
Siegel, Noah Y.
Perez-Ortiz, Maria
Treleaven, Philip
contents Large language models (LLMs) are becoming bigger to boost performance. However, little is known about how explainability is affected by this trend. This work explores LIME explanations for DeBERTaV3 models of four different sizes on natural language inference (NLI) and zero-shot classification (ZSC) tasks. We evaluate the explanations based on their faithfulness to the models' internal decision processes and their plausibility, i.e. their agreement with human explanations. The key finding is that increased model size does not correlate with plausibility despite improved model performance, suggesting a misalignment between the LIME explanations and the models' internal processes as model size increases. Our results further suggest limitations regarding faithfulness metrics in NLI contexts.
format Preprint
id arxiv_https___arxiv_org_abs_2405_05348
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle The Effect of Model Size on LLM Post-hoc Explainability via LIME
Heyen, Henning
Widdicombe, Amy
Siegel, Noah Y.
Perez-Ortiz, Maria
Treleaven, Philip
Computation and Language
Artificial Intelligence
Machine Learning
Large language models (LLMs) are becoming bigger to boost performance. However, little is known about how explainability is affected by this trend. This work explores LIME explanations for DeBERTaV3 models of four different sizes on natural language inference (NLI) and zero-shot classification (ZSC) tasks. We evaluate the explanations based on their faithfulness to the models' internal decision processes and their plausibility, i.e. their agreement with human explanations. The key finding is that increased model size does not correlate with plausibility despite improved model performance, suggesting a misalignment between the LIME explanations and the models' internal processes as model size increases. Our results further suggest limitations regarding faithfulness metrics in NLI contexts.
title The Effect of Model Size on LLM Post-hoc Explainability via LIME
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2405.05348