Linguistic Fingerprint in Transformer Models: How Language Variation Influences Parameter Selection in Irony Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mastromattei, Michele, Zanzotto, Fabio Massimo
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913377033388032
author Mastromattei, Michele
Zanzotto, Fabio Massimo
author_facet Mastromattei, Michele
Zanzotto, Fabio Massimo
contents This paper explores the correlation between linguistic diversity, sentiment analysis and transformer model architectures. We aim to investigate how different English variations impact transformer-based models for irony detection. To conduct our study, we used the EPIC corpus to extract five diverse English variation-specific datasets and applied the KEN pruning algorithm on five different architectures. Our results reveal several similarities between optimal subnetworks, which provide insights into the linguistic variations that share strong resemblances and those that exhibit greater dissimilarities. We discovered that optimal subnetworks across models share at least 60% of their parameters, emphasizing the significance of parameter values in capturing and interpreting linguistic variations. This study highlights the inherent structural similarities between models trained on different variants of the same language and also the critical role of parameter values in capturing these nuances.
format Preprint
id arxiv_https___arxiv_org_abs_2406_02338
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Linguistic Fingerprint in Transformer Models: How Language Variation Influences Parameter Selection in Irony Detection
Mastromattei, Michele
Zanzotto, Fabio Massimo
Computation and Language
Artificial Intelligence
This paper explores the correlation between linguistic diversity, sentiment analysis and transformer model architectures. We aim to investigate how different English variations impact transformer-based models for irony detection. To conduct our study, we used the EPIC corpus to extract five diverse English variation-specific datasets and applied the KEN pruning algorithm on five different architectures. Our results reveal several similarities between optimal subnetworks, which provide insights into the linguistic variations that share strong resemblances and those that exhibit greater dissimilarities. We discovered that optimal subnetworks across models share at least 60% of their parameters, emphasizing the significance of parameter values in capturing and interpreting linguistic variations. This study highlights the inherent structural similarities between models trained on different variants of the same language and also the critical role of parameter values in capturing these nuances.
title Linguistic Fingerprint in Transformer Models: How Language Variation Influences Parameter Selection in Irony Detection
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2406.02338