Lla-VAP: LSTM Ensemble of Llama and VAP for Turn-Taking Prediction

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Jeon, Hyunbae, Guintu, Frederic, Sahni, Rayvant
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916541432332288
author Jeon, Hyunbae
Guintu, Frederic
Sahni, Rayvant
author_facet Jeon, Hyunbae
Guintu, Frederic
Sahni, Rayvant
contents Turn-taking prediction is the task of anticipating when the speaker in a conversation will yield their turn to another speaker to begin speaking. This project expands on existing strategies for turn-taking prediction by employing a multi-modal ensemble approach that integrates large language models (LLMs) and voice activity projection (VAP) models. By combining the linguistic capabilities of LLMs with the temporal precision of VAP models, we aim to improve the accuracy and efficiency of identifying TRPs in both scripted and unscripted conversational scenarios. Our methods are evaluated on the In-Conversation Corpus (ICC) and Coached Conversational Preference Elicitation (CCPE) datasets, highlighting the strengths and limitations of current models while proposing a potentially more robust framework for enhanced prediction.
format Preprint
id arxiv_https___arxiv_org_abs_2412_18061
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Lla-VAP: LSTM Ensemble of Llama and VAP for Turn-Taking Prediction
Jeon, Hyunbae
Guintu, Frederic
Sahni, Rayvant
Sound
Computation and Language
Human-Computer Interaction
Audio and Speech Processing
Turn-taking prediction is the task of anticipating when the speaker in a conversation will yield their turn to another speaker to begin speaking. This project expands on existing strategies for turn-taking prediction by employing a multi-modal ensemble approach that integrates large language models (LLMs) and voice activity projection (VAP) models. By combining the linguistic capabilities of LLMs with the temporal precision of VAP models, we aim to improve the accuracy and efficiency of identifying TRPs in both scripted and unscripted conversational scenarios. Our methods are evaluated on the In-Conversation Corpus (ICC) and Coached Conversational Preference Elicitation (CCPE) datasets, highlighting the strengths and limitations of current models while proposing a potentially more robust framework for enhanced prediction.
title Lla-VAP: LSTM Ensemble of Llama and VAP for Turn-Taking Prediction
topic Sound
Computation and Language
Human-Computer Interaction
Audio and Speech Processing
url https://arxiv.org/abs/2412.18061