Lla-VAP: LSTM Ensemble of Llama and VAP for Turn-Taking Prediction
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866916541432332288 |
|---|---|
| author | Jeon, Hyunbae Guintu, Frederic Sahni, Rayvant |
| author_facet | Jeon, Hyunbae Guintu, Frederic Sahni, Rayvant |
| contents | Turn-taking prediction is the task of anticipating when the speaker in a conversation will yield their turn to another speaker to begin speaking. This project expands on existing strategies for turn-taking prediction by employing a multi-modal ensemble approach that integrates large language models (LLMs) and voice activity projection (VAP) models. By combining the linguistic capabilities of LLMs with the temporal precision of VAP models, we aim to improve the accuracy and efficiency of identifying TRPs in both scripted and unscripted conversational scenarios. Our methods are evaluated on the In-Conversation Corpus (ICC) and Coached Conversational Preference Elicitation (CCPE) datasets, highlighting the strengths and limitations of current models while proposing a potentially more robust framework for enhanced prediction. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_18061 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Lla-VAP: LSTM Ensemble of Llama and VAP for Turn-Taking Prediction Jeon, Hyunbae Guintu, Frederic Sahni, Rayvant Sound Computation and Language Human-Computer Interaction Audio and Speech Processing Turn-taking prediction is the task of anticipating when the speaker in a conversation will yield their turn to another speaker to begin speaking. This project expands on existing strategies for turn-taking prediction by employing a multi-modal ensemble approach that integrates large language models (LLMs) and voice activity projection (VAP) models. By combining the linguistic capabilities of LLMs with the temporal precision of VAP models, we aim to improve the accuracy and efficiency of identifying TRPs in both scripted and unscripted conversational scenarios. Our methods are evaluated on the In-Conversation Corpus (ICC) and Coached Conversational Preference Elicitation (CCPE) datasets, highlighting the strengths and limitations of current models while proposing a potentially more robust framework for enhanced prediction. |
| title | Lla-VAP: LSTM Ensemble of Llama and VAP for Turn-Taking Prediction |
| topic | Sound Computation and Language Human-Computer Interaction Audio and Speech Processing |
| url | https://arxiv.org/abs/2412.18061 |