Lla-VAP: LSTM Ensemble of Llama and VAP for Turn-Taking Prediction
Fuente:
arXiv
Saved in:
| Main Authors: | Jeon, Hyunbae, Guintu, Frederic, Sahni, Rayvant |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Real-time and Continuous Turn-taking Prediction Using Voice Activity Projection
by: Inoue, Koji, et al.
Published: (2024)
by: Inoue, Koji, et al.
Published: (2024)
Yeah, Un, Oh: Continuous and Real-time Backchannel Prediction with Fine-tuning of Voice Activity Projection
by: Inoue, Koji, et al.
Published: (2024)
by: Inoue, Koji, et al.
Published: (2024)
Early Detection of Furniture-Infesting Wood-Boring Beetles Using CNN-LSTM Networks and MFCC-Based Acoustic Features
by: Manukalpa, J. M. Chan Sri, et al.
Published: (2025)
by: Manukalpa, J. M. Chan Sri, et al.
Published: (2025)
Speech vs. Transcript: Does It Matter for Human Annotators in Speech Summarization?
by: Sharma, Roshan, et al.
Published: (2024)
by: Sharma, Roshan, et al.
Published: (2024)
Enhancing AAC Software for Dysarthric Speakers in e-Health Settings: An Evaluation Using TORGO
by: Hui, Macarious, et al.
Published: (2024)
by: Hui, Macarious, et al.
Published: (2024)
MCMChaos: Improvising Rap Music with MCMC Methods and Chaos Theory
by: Kimelman, Robert G.
Published: (2024)
by: Kimelman, Robert G.
Published: (2024)
Detecting the terminality of speech-turn boundary for spoken interactions in French TV and Radio content
by: Uro, Rémi, et al.
Published: (2024)
by: Uro, Rémi, et al.
Published: (2024)
Investigating the Effects of Large-Scale Pseudo-Stereo Data and Different Speech Foundation Model on Dialogue Generative Spoken Language Model
by: Fu, Yu-Kuan, et al.
Published: (2024)
by: Fu, Yu-Kuan, et al.
Published: (2024)
OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios
by: Cheng, Xize, et al.
Published: (2025)
by: Cheng, Xize, et al.
Published: (2025)
The ICASSP 2026 HumDial Challenge: Benchmarking Human-like Spoken Dialogue Systems in the LLM Era
by: Zhao, Zhixian, et al.
Published: (2026)
by: Zhao, Zhixian, et al.
Published: (2026)
Are Expressions for Music Emotions the Same Across Cultures?
by: Celen, Elif, et al.
Published: (2025)
by: Celen, Elif, et al.
Published: (2025)
InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-training
by: Wang, Dingdong, et al.
Published: (2025)
by: Wang, Dingdong, et al.
Published: (2025)
Loop Copilot: Conducting AI Ensembles for Music Generation and Iterative Editing
by: Zhang, Yixiao, et al.
Published: (2023)
by: Zhang, Yixiao, et al.
Published: (2023)
Efficient Ensemble for Multimodal Punctuation Restoration using Time-Delay Neural Network
by: Liu, Xing Yi, et al.
Published: (2023)
by: Liu, Xing Yi, et al.
Published: (2023)
Prompt-Guided Turn-Taking Prediction
by: Inoue, Koji, et al.
Published: (2025)
by: Inoue, Koji, et al.
Published: (2025)
Open Your Ears and Take a Look: A State-of-the-Art Report on the Integration of Sonification and Visualization
by: Enge, Kajetan, et al.
Published: (2024)
by: Enge, Kajetan, et al.
Published: (2024)
LSTM-CNN Network for Audio Signature Analysis in Noisy Environments
by: Damacharla, Praveen, et al.
Published: (2023)
by: Damacharla, Praveen, et al.
Published: (2023)
Towards Reliable Large Audio Language Model
by: Ma, Ziyang, et al.
Published: (2025)
by: Ma, Ziyang, et al.
Published: (2025)
A Mapping Strategy for Interacting with Latent Audio Synthesis Using Artistic Materials
by: Zheng, Shuoyang, et al.
Published: (2024)
by: Zheng, Shuoyang, et al.
Published: (2024)
Enhancing DMI Interactions by Integrating Haptic Feedback for Intricate Vibrato Technique
by: Piao, Ziyue, et al.
Published: (2024)
by: Piao, Ziyue, et al.
Published: (2024)
A cross-talk robust multichannel VAD model for multiparty agent interactions trained using synthetic re-recordings
by: Han, Hyewon, et al.
Published: (2024)
by: Han, Hyewon, et al.
Published: (2024)
Interactive Sonification for Health and Energy using ChucK and Unity
by: Zhao, Yichun, et al.
Published: (2024)
by: Zhao, Yichun, et al.
Published: (2024)
A Near-Real-Time Processing Ego Speech Filtering Pipeline Designed for Speech Interruption During Human-Robot Interaction
by: Li, Yue, et al.
Published: (2024)
by: Li, Yue, et al.
Published: (2024)
USpeech: Ultrasound-Enhanced Speech with Minimal Human Effort via Cross-Modal Synthesis
by: Yu, Luca Jiang-Tao, et al.
Published: (2024)
by: Yu, Luca Jiang-Tao, et al.
Published: (2024)
Interfacing with history: Curating with audio augmented objects
by: Cliffe, Laurence
Published: (2024)
by: Cliffe, Laurence
Published: (2024)
Transhuman Ansambl - Voice Beyond Language
by: Ivsic, Lucija, et al.
Published: (2024)
by: Ivsic, Lucija, et al.
Published: (2024)
How Private is Low-Frequency Speech Audio in the Wild? An Analysis of Verbal Intelligibility by Humans and Machines
by: Liu, Ailin, et al.
Published: (2024)
by: Liu, Ailin, et al.
Published: (2024)
Cervical Auscultation Machine Learning for Dysphagia Assessment
by: Chia, An An, et al.
Published: (2024)
by: Chia, An An, et al.
Published: (2024)
Springboard, Roadblock or "Crutch"?: How Transgender Users Leverage Voice Changers for Gender Presentation in Social Virtual Reality
by: Povinelli, Kassie, et al.
Published: (2024)
by: Povinelli, Kassie, et al.
Published: (2024)
Teach Me How to ImproVISe: Co-Designing an Augmented Piano Training System for Improvisation
by: Deja, Jordan Aiko, et al.
Published: (2024)
by: Deja, Jordan Aiko, et al.
Published: (2024)
Open vocabulary keyword spotting through transfer learning from speech synthesis
by: V, Kesavaraj, et al.
Published: (2024)
by: V, Kesavaraj, et al.
Published: (2024)
The effect of self-motion and room familiarity on sound source localization in virtual environments
by: Isserstedt, Niklas, et al.
Published: (2024)
by: Isserstedt, Niklas, et al.
Published: (2024)
NeckCare: Preventing Tech Neck using Hearable-based Multimodal Sensing
by: Chhaglani, Bhawana, et al.
Published: (2024)
by: Chhaglani, Bhawana, et al.
Published: (2024)
Optimizing Dysarthria Wake-Up Word Spotting: An End-to-End Approach for SLT 2024 LRDWWS Challenge
by: Liu, Shuiyun, et al.
Published: (2024)
by: Liu, Shuiyun, et al.
Published: (2024)
Collaboration Between Robots, Interfaces and Humans: Practice-Based and Audience Perspectives
by: Savery, Anna, et al.
Published: (2024)
by: Savery, Anna, et al.
Published: (2024)
A Framework for AI assisted Musical Devices
by: Civit, Miguel, et al.
Published: (2024)
by: Civit, Miguel, et al.
Published: (2024)
RespEar: Earable-Based Robust Respiratory Rate Monitoring
by: Liu, Yang, et al.
Published: (2024)
by: Liu, Yang, et al.
Published: (2024)
Using Confidence Scores to Improve Eyes-free Detection of Speech Recognition Errors
by: Nowrin, Sadia, et al.
Published: (2024)
by: Nowrin, Sadia, et al.
Published: (2024)
SoundShift: Exploring Sound Manipulations for Accessible Mixed-Reality Awareness
by: Chang, Ruei-Che, et al.
Published: (2024)
by: Chang, Ruei-Che, et al.
Published: (2024)
Improving Multimodal Emotion Recognition by Leveraging Acoustic Adaptation and Visual Alignment
by: Zhao, Zhixian, et al.
Published: (2024)
by: Zhao, Zhixian, et al.
Published: (2024)
Similar Items
-
Real-time and Continuous Turn-taking Prediction Using Voice Activity Projection
by: Inoue, Koji, et al.
Published: (2024) -
Yeah, Un, Oh: Continuous and Real-time Backchannel Prediction with Fine-tuning of Voice Activity Projection
by: Inoue, Koji, et al.
Published: (2024) -
Early Detection of Furniture-Infesting Wood-Boring Beetles Using CNN-LSTM Networks and MFCC-Based Acoustic Features
by: Manukalpa, J. M. Chan Sri, et al.
Published: (2025) -
Speech vs. Transcript: Does It Matter for Human Annotators in Speech Summarization?
by: Sharma, Roshan, et al.
Published: (2024) -
Enhancing AAC Software for Dysarthric Speakers in e-Health Settings: An Evaluation Using TORGO
by: Hui, Macarious, et al.
Published: (2024)