Baby Bear: Seeking a Just Right Rating Scale for Scalar Annotations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Han, Xu, Yu, Felix, Sedoc, Joao, Van Durme, Benjamin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910570408574976
author Han, Xu
Yu, Felix
Sedoc, Joao
Van Durme, Benjamin
author_facet Han, Xu
Yu, Felix
Sedoc, Joao
Van Durme, Benjamin
contents Our goal is a mechanism for efficiently assigning scalar ratings to each of a large set of elements. For example, "what percent positive or negative is this product review?" When sample sizes are small, prior work has advocated for methods such as Best Worst Scaling (BWS) as being more robust than direct ordinal annotation ("Likert scales"). Here we first introduce IBWS, which iteratively collects annotations through Best-Worst Scaling, resulting in robustly ranked crowd-sourced data. While effective, IBWS is too expensive for large-scale tasks. Using the results of IBWS as a best-desired outcome, we evaluate various direct assessment methods to determine what is both cost-efficient and best correlating to a large scale BWS annotation strategy. Finally, we illustrate in the domains of dialogue and sentiment how these annotations can support robust learning-to-rank models.
format Preprint
id arxiv_https___arxiv_org_abs_2408_09765
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Baby Bear: Seeking a Just Right Rating Scale for Scalar Annotations
Han, Xu
Yu, Felix
Sedoc, Joao
Van Durme, Benjamin
Machine Learning
Human-Computer Interaction
Our goal is a mechanism for efficiently assigning scalar ratings to each of a large set of elements. For example, "what percent positive or negative is this product review?" When sample sizes are small, prior work has advocated for methods such as Best Worst Scaling (BWS) as being more robust than direct ordinal annotation ("Likert scales"). Here we first introduce IBWS, which iteratively collects annotations through Best-Worst Scaling, resulting in robustly ranked crowd-sourced data. While effective, IBWS is too expensive for large-scale tasks. Using the results of IBWS as a best-desired outcome, we evaluate various direct assessment methods to determine what is both cost-efficient and best correlating to a large scale BWS annotation strategy. Finally, we illustrate in the domains of dialogue and sentiment how these annotations can support robust learning-to-rank models.
title Baby Bear: Seeking a Just Right Rating Scale for Scalar Annotations
topic Machine Learning
Human-Computer Interaction
url https://arxiv.org/abs/2408.09765