ImitAL: Learned Active Learning Strategy on Synthetic Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gonsior, Julius, Thiele, Maik, Lehner, Wolfgang
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910870547726336
author Gonsior, Julius
Thiele, Maik
Lehner, Wolfgang
author_facet Gonsior, Julius
Thiele, Maik
Lehner, Wolfgang
contents Active Learning (AL) is a well-known standard method for efficiently obtaining annotated data by first labeling the samples that contain the most information based on a query strategy. In the past, a large variety of such query strategies has been proposed, with each generation of new strategies increasing the runtime and adding more complexity. However, to the best of our our knowledge, none of these strategies excels consistently over a large number of datasets from different application domains. Basically, most of the the existing AL strategies are a combination of the two simple heuristics informativeness and representativeness, and the big differences lie in the combination of the often conflicting heuristics. Within this paper, we propose ImitAL, a domain-independent novel query strategy, which encodes AL as a learning-to-rank problem and learns an optimal combination between both heuristics. We train ImitAL on large-scale simulated AL runs on purely synthetic datasets. To show that ImitAL was successfully trained, we perform an extensive evaluation comparing our strategy on 13 different datasets, from a wide range of domains, with 7 other query strategies.
format Preprint
id arxiv_https___arxiv_org_abs_2208_11636
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle ImitAL: Learned Active Learning Strategy on Synthetic Data
Gonsior, Julius
Thiele, Maik
Lehner, Wolfgang
Machine Learning
Artificial Intelligence
Active Learning (AL) is a well-known standard method for efficiently obtaining annotated data by first labeling the samples that contain the most information based on a query strategy. In the past, a large variety of such query strategies has been proposed, with each generation of new strategies increasing the runtime and adding more complexity. However, to the best of our our knowledge, none of these strategies excels consistently over a large number of datasets from different application domains. Basically, most of the the existing AL strategies are a combination of the two simple heuristics informativeness and representativeness, and the big differences lie in the combination of the often conflicting heuristics. Within this paper, we propose ImitAL, a domain-independent novel query strategy, which encodes AL as a learning-to-rank problem and learns an optimal combination between both heuristics. We train ImitAL on large-scale simulated AL runs on purely synthetic datasets. To show that ImitAL was successfully trained, we perform an extensive evaluation comparing our strategy on 13 different datasets, from a wide range of domains, with 7 other query strategies.
title ImitAL: Learned Active Learning Strategy on Synthetic Data
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2208.11636