Using Shapley interactions to understand how models use structure

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Singhvi, Divyansh, Misra, Diganta, Erkelens, Andrej, Jain, Raghav, Papadimitriou, Isabel, Saphra, Naomi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916788861665280
author Singhvi, Divyansh
Misra, Diganta
Erkelens, Andrej
Jain, Raghav
Papadimitriou, Isabel
Saphra, Naomi
author_facet Singhvi, Divyansh
Misra, Diganta
Erkelens, Andrej
Jain, Raghav
Papadimitriou, Isabel
Saphra, Naomi
contents Language is an intricately structured system, and a key goal of NLP interpretability is to provide methodological insights for understanding how language models represent this structure internally. In this paper, we use Shapley Taylor interaction indices (STII) in order to examine how language and speech models internally relate and structure their inputs. Pairwise Shapley interactions measure how much two inputs work together to influence model outputs beyond if we linearly added their independent influences, providing a view into how models encode structural interactions between inputs. We relate the interaction patterns in models to three underlying linguistic structures: syntactic structure, non-compositional semantics, and phonetic coarticulation. We find that autoregressive text models encode interactions that correlate with the syntactic proximity of inputs, and that both autoregressive and masked models encode nonlinear interactions in idiomatic phrases with non-compositional semantics. Our speech results show that inputs are more entangled for pairs where a neighboring consonant is likely to influence a vowel or approximant, showing that models encode the phonetic interaction needed for extracting discrete phonemic representations.
format Preprint
id arxiv_https___arxiv_org_abs_2403_13106
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Using Shapley interactions to understand how models use structure
Singhvi, Divyansh
Misra, Diganta
Erkelens, Andrej
Jain, Raghav
Papadimitriou, Isabel
Saphra, Naomi
Machine Learning
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Language is an intricately structured system, and a key goal of NLP interpretability is to provide methodological insights for understanding how language models represent this structure internally. In this paper, we use Shapley Taylor interaction indices (STII) in order to examine how language and speech models internally relate and structure their inputs. Pairwise Shapley interactions measure how much two inputs work together to influence model outputs beyond if we linearly added their independent influences, providing a view into how models encode structural interactions between inputs. We relate the interaction patterns in models to three underlying linguistic structures: syntactic structure, non-compositional semantics, and phonetic coarticulation. We find that autoregressive text models encode interactions that correlate with the syntactic proximity of inputs, and that both autoregressive and masked models encode nonlinear interactions in idiomatic phrases with non-compositional semantics. Our speech results show that inputs are more entangled for pairs where a neighboring consonant is likely to influence a vowel or approximant, showing that models encode the phonetic interaction needed for extracting discrete phonemic representations.
title Using Shapley interactions to understand how models use structure
topic Machine Learning
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.13106