Saved in:
Bibliographic Details
Main Authors: Azeemi, Abdul Hameed, Qazi, Ihsan Ayyub, Raza, Agha Ali
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2403.09259
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929637065490432
author Azeemi, Abdul Hameed
Qazi, Ihsan Ayyub
Raza, Agha Ali
author_facet Azeemi, Abdul Hameed
Qazi, Ihsan Ayyub
Raza, Agha Ali
contents Active learning (AL) techniques reduce labeling costs for training neural machine translation (NMT) models by selecting smaller representative subsets from unlabeled data for annotation. Diversity sampling techniques select heterogeneous instances, while uncertainty sampling methods select instances with the highest model uncertainty. Both approaches have limitations - diversity methods may extract varied but trivial examples, while uncertainty sampling can yield repetitive, uninformative instances. To bridge this gap, we propose Hybrid Uncertainty and Diversity Sampling (HUDS), an AL strategy for domain adaptation in NMT that combines uncertainty and diversity for sentence selection. HUDS computes uncertainty scores for unlabeled sentences and subsequently stratifies them. It then clusters sentence embeddings within each stratum and computes diversity scores by distance to the centroid. A weighted hybrid score that combines uncertainty and diversity is then used to select the top instances for annotation in each AL iteration. Experiments on multi-domain German-English and French-English datasets demonstrate the better performance of HUDS over other strong AL baselines. We analyze the sentence selection with HUDS and show that it prioritizes diverse instances having high model uncertainty for annotation in early AL iterations.
format Preprint
id arxiv_https___arxiv_org_abs_2403_09259
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle To Label or Not to Label: Hybrid Active Learning for Neural Machine Translation
Azeemi, Abdul Hameed
Qazi, Ihsan Ayyub
Raza, Agha Ali
Computation and Language
Machine Learning
Active learning (AL) techniques reduce labeling costs for training neural machine translation (NMT) models by selecting smaller representative subsets from unlabeled data for annotation. Diversity sampling techniques select heterogeneous instances, while uncertainty sampling methods select instances with the highest model uncertainty. Both approaches have limitations - diversity methods may extract varied but trivial examples, while uncertainty sampling can yield repetitive, uninformative instances. To bridge this gap, we propose Hybrid Uncertainty and Diversity Sampling (HUDS), an AL strategy for domain adaptation in NMT that combines uncertainty and diversity for sentence selection. HUDS computes uncertainty scores for unlabeled sentences and subsequently stratifies them. It then clusters sentence embeddings within each stratum and computes diversity scores by distance to the centroid. A weighted hybrid score that combines uncertainty and diversity is then used to select the top instances for annotation in each AL iteration. Experiments on multi-domain German-English and French-English datasets demonstrate the better performance of HUDS over other strong AL baselines. We analyze the sentence selection with HUDS and show that it prioritizes diverse instances having high model uncertainty for annotation in early AL iterations.
title To Label or Not to Label: Hybrid Active Learning for Neural Machine Translation
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2403.09259