Beyond Words: Interjection Classification for Improved Human-Computer Interaction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Goren, Yaniv, Cohen, Yuval, Apartsin, Alexander, Aperstein, Yehudit
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918135013048320
author Goren, Yaniv
Cohen, Yuval
Apartsin, Alexander
Aperstein, Yehudit
author_facet Goren, Yaniv
Cohen, Yuval
Apartsin, Alexander
Aperstein, Yehudit
contents In the realm of human-computer interaction, fostering a natural dialogue between humans and machines is paramount. A key, often overlooked, component of this dialogue is the use of interjections such as "mmm" and "hmm". Despite their frequent use to express agreement, hesitation, or requests for information, these interjections are typically dismissed as "non-words" by Automatic Speech Recognition (ASR) engines. Addressing this gap, we introduce a novel task dedicated to interjection classification, a pioneer in the field to our knowledge. This task is challenging due to the short duration of interjection signals and significant inter- and intra-speaker variability. In this work, we present and publish a dataset of interjection signals collected specifically for interjection classification. We employ this dataset to train and evaluate a baseline deep learning model. To enhance performance, we augment the training dataset using techniques such as tempo and pitch transformation, which significantly improve classification accuracy, making models more robust. The interjection dataset, a Python library for the augmentation pipeline, baseline model, and evaluation scripts, are available to the research community.
format Preprint
id arxiv_https___arxiv_org_abs_2509_03181
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond Words: Interjection Classification for Improved Human-Computer Interaction
Goren, Yaniv
Cohen, Yuval
Apartsin, Alexander
Aperstein, Yehudit
Human-Computer Interaction
Machine Learning
In the realm of human-computer interaction, fostering a natural dialogue between humans and machines is paramount. A key, often overlooked, component of this dialogue is the use of interjections such as "mmm" and "hmm". Despite their frequent use to express agreement, hesitation, or requests for information, these interjections are typically dismissed as "non-words" by Automatic Speech Recognition (ASR) engines. Addressing this gap, we introduce a novel task dedicated to interjection classification, a pioneer in the field to our knowledge. This task is challenging due to the short duration of interjection signals and significant inter- and intra-speaker variability. In this work, we present and publish a dataset of interjection signals collected specifically for interjection classification. We employ this dataset to train and evaluate a baseline deep learning model. To enhance performance, we augment the training dataset using techniques such as tempo and pitch transformation, which significantly improve classification accuracy, making models more robust. The interjection dataset, a Python library for the augmentation pipeline, baseline model, and evaluation scripts, are available to the research community.
title Beyond Words: Interjection Classification for Improved Human-Computer Interaction
topic Human-Computer Interaction
Machine Learning
url https://arxiv.org/abs/2509.03181