No More Distractions: an Adaptive Up-Sampling Algorithm to Reduce Data Artifacts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Chen, Han
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910307765452800
author Chen, Han
author_facet Chen, Han
contents Researchers recently found out that sometimes language models achieve high accuracy on benchmark data set, but they can not generalize very well with even little changes to the original data set. This is sometimes due to data artifacts, model is learning the spurious correlation between tokens and labels, instead of the semantics and logic. In this work, we analyzed SNLI data and visualized such spurious correlations. We proposed an adaptive up-sampling algorithm to correct the data artifacts, which is simple and effective, and does not need human edits or annotation. We did an experiment applying the algorithm to fix the data artifacts in SNLI data and the model trained with corrected data performed significantly better than the model trained with raw SNLI data, overall, as well as on the subset we corrected.
format Preprint
id arxiv_https___arxiv_org_abs_2401_13907
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle No More Distractions: an Adaptive Up-Sampling Algorithm to Reduce Data Artifacts
Chen, Han
Computation and Language
Researchers recently found out that sometimes language models achieve high accuracy on benchmark data set, but they can not generalize very well with even little changes to the original data set. This is sometimes due to data artifacts, model is learning the spurious correlation between tokens and labels, instead of the semantics and logic. In this work, we analyzed SNLI data and visualized such spurious correlations. We proposed an adaptive up-sampling algorithm to correct the data artifacts, which is simple and effective, and does not need human edits or annotation. We did an experiment applying the algorithm to fix the data artifacts in SNLI data and the model trained with corrected data performed significantly better than the model trained with raw SNLI data, overall, as well as on the subset we corrected.
title No More Distractions: an Adaptive Up-Sampling Algorithm to Reduce Data Artifacts
topic Computation and Language
url https://arxiv.org/abs/2401.13907