Local Interpretations for Explainable Natural Language Processing: A Survey

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Luo, Siwen, Ivison, Hamish, Han, Caren, Poon, Josiah
Natura: Preprint
Pubblicazione: 2021
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909138499403776
author Luo, Siwen
Ivison, Hamish
Han, Caren
Poon, Josiah
author_facet Luo, Siwen
Ivison, Hamish
Han, Caren
Poon, Josiah
contents As the use of deep learning techniques has grown across various fields over the past decade, complaints about the opaqueness of the black-box models have increased, resulting in an increased focus on transparency in deep learning models. This work investigates various methods to improve the interpretability of deep neural networks for Natural Language Processing (NLP) tasks, including machine translation and sentiment analysis. We provide a comprehensive discussion on the definition of the term interpretability and its various aspects at the beginning of this work. The methods collected and summarised in this survey are only associated with local interpretation and are specifically divided into three categories: 1) interpreting the model's predictions through related input features; 2) interpreting through natural language explanation; 3) probing the hidden states of models and word representations.
format Preprint
id arxiv_https___arxiv_org_abs_2103_11072
institution arXiv
publishDate 2021
record_format arxiv
spellingShingle Local Interpretations for Explainable Natural Language Processing: A Survey
Luo, Siwen
Ivison, Hamish
Han, Caren
Poon, Josiah
Computation and Language
Artificial Intelligence
A.1; I.2.7
As the use of deep learning techniques has grown across various fields over the past decade, complaints about the opaqueness of the black-box models have increased, resulting in an increased focus on transparency in deep learning models. This work investigates various methods to improve the interpretability of deep neural networks for Natural Language Processing (NLP) tasks, including machine translation and sentiment analysis. We provide a comprehensive discussion on the definition of the term interpretability and its various aspects at the beginning of this work. The methods collected and summarised in this survey are only associated with local interpretation and are specifically divided into three categories: 1) interpreting the model's predictions through related input features; 2) interpreting through natural language explanation; 3) probing the hidden states of models and word representations.
title Local Interpretations for Explainable Natural Language Processing: A Survey
topic Computation and Language
Artificial Intelligence
A.1; I.2.7
url https://arxiv.org/abs/2103.11072