Efficient Ensemble for Multimodal Punctuation Restoration using Time-Delay Neural Network
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Xing Yi, Beigi, Homayoon |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Spontaneous Informal Speech Dataset for Punctuation Restoration
von: Liu, Xing Yi, et al.
Veröffentlicht: (2024)
von: Liu, Xing Yi, et al.
Veröffentlicht: (2024)
Carnatic Raga Identification System using Rigorous Time-Delay Neural Network
von: Natesan, Sanjay, et al.
Veröffentlicht: (2024)
von: Natesan, Sanjay, et al.
Veröffentlicht: (2024)
Lla-VAP: LSTM Ensemble of Llama and VAP for Turn-Taking Prediction
von: Jeon, Hyunbae, et al.
Veröffentlicht: (2024)
von: Jeon, Hyunbae, et al.
Veröffentlicht: (2024)
Loop Copilot: Conducting AI Ensembles for Music Generation and Iterative Editing
von: Zhang, Yixiao, et al.
Veröffentlicht: (2023)
von: Zhang, Yixiao, et al.
Veröffentlicht: (2023)
NeckCare: Preventing Tech Neck using Hearable-based Multimodal Sensing
von: Chhaglani, Bhawana, et al.
Veröffentlicht: (2024)
von: Chhaglani, Bhawana, et al.
Veröffentlicht: (2024)
LLAMAPIE: Proactive In-Ear Conversation Assistants
von: Chen, Tuochao, et al.
Veröffentlicht: (2025)
von: Chen, Tuochao, et al.
Veröffentlicht: (2025)
Artificial Neural Networks to Recognize Speakers Division from Continuous Bengali Speech
von: Ali, Hasmot, et al.
Veröffentlicht: (2024)
von: Ali, Hasmot, et al.
Veröffentlicht: (2024)
The ICASSP 2026 HumDial Challenge: Benchmarking Human-like Spoken Dialogue Systems in the LLM Era
von: Zhao, Zhixian, et al.
Veröffentlicht: (2026)
von: Zhao, Zhixian, et al.
Veröffentlicht: (2026)
Real-time and Continuous Turn-taking Prediction Using Voice Activity Projection
von: Inoue, Koji, et al.
Veröffentlicht: (2024)
von: Inoue, Koji, et al.
Veröffentlicht: (2024)
Yeah, Un, Oh: Continuous and Real-time Backchannel Prediction with Fine-tuning of Voice Activity Projection
von: Inoue, Koji, et al.
Veröffentlicht: (2024)
von: Inoue, Koji, et al.
Veröffentlicht: (2024)
OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios
von: Cheng, Xize, et al.
Veröffentlicht: (2025)
von: Cheng, Xize, et al.
Veröffentlicht: (2025)
Are Expressions for Music Emotions the Same Across Cultures?
von: Celen, Elif, et al.
Veröffentlicht: (2025)
von: Celen, Elif, et al.
Veröffentlicht: (2025)
Speech vs. Transcript: Does It Matter for Human Annotators in Speech Summarization?
von: Sharma, Roshan, et al.
Veröffentlicht: (2024)
von: Sharma, Roshan, et al.
Veröffentlicht: (2024)
Enhancing AAC Software for Dysarthric Speakers in e-Health Settings: An Evaluation Using TORGO
von: Hui, Macarious, et al.
Veröffentlicht: (2024)
von: Hui, Macarious, et al.
Veröffentlicht: (2024)
InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-training
von: Wang, Dingdong, et al.
Veröffentlicht: (2025)
von: Wang, Dingdong, et al.
Veröffentlicht: (2025)
MCMChaos: Improvising Rap Music with MCMC Methods and Chaos Theory
von: Kimelman, Robert G.
Veröffentlicht: (2024)
von: Kimelman, Robert G.
Veröffentlicht: (2024)
Detecting the terminality of speech-turn boundary for spoken interactions in French TV and Radio content
von: Uro, Rémi, et al.
Veröffentlicht: (2024)
von: Uro, Rémi, et al.
Veröffentlicht: (2024)
Investigating the Effects of Large-Scale Pseudo-Stereo Data and Different Speech Foundation Model on Dialogue Generative Spoken Language Model
von: Fu, Yu-Kuan, et al.
Veröffentlicht: (2024)
von: Fu, Yu-Kuan, et al.
Veröffentlicht: (2024)
VoXtream: Full-Stream Text-to-Speech with Extremely Low Latency
von: Torgashov, Nikita, et al.
Veröffentlicht: (2025)
von: Torgashov, Nikita, et al.
Veröffentlicht: (2025)
Abjad-Kids: An Arabic Speech Classification Dataset for Primary Education
von: Snoubara, Abdul Aziz, et al.
Veröffentlicht: (2026)
von: Snoubara, Abdul Aziz, et al.
Veröffentlicht: (2026)
AIx Speed: Playback Speed Optimization Using Listening Comprehension of Speech Recognition Models
von: Kawamura, Kazuki, et al.
Veröffentlicht: (2024)
von: Kawamura, Kazuki, et al.
Veröffentlicht: (2024)
A conversational gesture synthesis system based on emotions and semantics
von: Hoang-Minh, Thanh
Veröffentlicht: (2025)
von: Hoang-Minh, Thanh
Veröffentlicht: (2025)
VoXtream2: Full-stream TTS with dynamic speaking rate control
von: Torgashov, Nikita, et al.
Veröffentlicht: (2026)
von: Torgashov, Nikita, et al.
Veröffentlicht: (2026)
Literary and Colloquial Tamil Dialect Identification
von: Nanmalar, M., et al.
Veröffentlicht: (2024)
von: Nanmalar, M., et al.
Veröffentlicht: (2024)
Improving Multimodal Emotion Recognition by Leveraging Acoustic Adaptation and Visual Alignment
von: Zhao, Zhixian, et al.
Veröffentlicht: (2024)
von: Zhao, Zhixian, et al.
Veröffentlicht: (2024)
Subject Disentanglement Neural Network for Speech Envelope Reconstruction from EEG
von: Zhang, Li, et al.
Veröffentlicht: (2025)
von: Zhang, Li, et al.
Veröffentlicht: (2025)
NeuroIncept Decoder for High-Fidelity Speech Reconstruction from Neural Activity
von: Khanday, Owais Mujtaba, et al.
Veröffentlicht: (2025)
von: Khanday, Owais Mujtaba, et al.
Veröffentlicht: (2025)
Recreating Neural Activity During Speech Production with Language and Speech Model Embeddings
von: Khanday, Owais Mujtaba, et al.
Veröffentlicht: (2025)
von: Khanday, Owais Mujtaba, et al.
Veröffentlicht: (2025)
Adapting Whisper for Lightweight and Efficient Automatic Speech Recognition of Children for On-device Edge Applications
von: Dutta, Satwik, et al.
Veröffentlicht: (2025)
von: Dutta, Satwik, et al.
Veröffentlicht: (2025)
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
von: Chen, Qian, et al.
Veröffentlicht: (2025)
von: Chen, Qian, et al.
Veröffentlicht: (2025)
MHANet: Multi-scale Hybrid Attention Network for Auditory Attention Detection
von: Li, Lu, et al.
Veröffentlicht: (2025)
von: Li, Lu, et al.
Veröffentlicht: (2025)
Psychophysiology-aided Perceptually Fluent Speech Analysis of Children Who Stutter
von: Xiao, Yi, et al.
Veröffentlicht: (2022)
von: Xiao, Yi, et al.
Veröffentlicht: (2022)
Interactive Sonification for Health and Energy using ChucK and Unity
von: Zhao, Yichun, et al.
Veröffentlicht: (2024)
von: Zhao, Yichun, et al.
Veröffentlicht: (2024)
A Near-Real-Time Processing Ego Speech Filtering Pipeline Designed for Speech Interruption During Human-Robot Interaction
von: Li, Yue, et al.
Veröffentlicht: (2024)
von: Li, Yue, et al.
Veröffentlicht: (2024)
ListenNet: A Lightweight Spatio-Temporal Enhancement Nested Network for Auditory Attention Detection
von: Fan, Cunhang, et al.
Veröffentlicht: (2025)
von: Fan, Cunhang, et al.
Veröffentlicht: (2025)
Early Detection of Furniture-Infesting Wood-Boring Beetles Using CNN-LSTM Networks and MFCC-Based Acoustic Features
von: Manukalpa, J. M. Chan Sri, et al.
Veröffentlicht: (2025)
von: Manukalpa, J. M. Chan Sri, et al.
Veröffentlicht: (2025)
A cross-talk robust multichannel VAD model for multiparty agent interactions trained using synthetic re-recordings
von: Han, Hyewon, et al.
Veröffentlicht: (2024)
von: Han, Hyewon, et al.
Veröffentlicht: (2024)
AAD-LLM: Neural Attention-Driven Auditory Scene Understanding
von: Jiang, Xilin, et al.
Veröffentlicht: (2025)
von: Jiang, Xilin, et al.
Veröffentlicht: (2025)
Beyond Acoustic Emotion Recognition: Multimodal Pathos Analysis in Political Speech Using LLM-Based and Acoustic Emotion Models
von: Dietrich, Juergen
Veröffentlicht: (2026)
von: Dietrich, Juergen
Veröffentlicht: (2026)
Towards Reliable Large Audio Language Model
von: Ma, Ziyang, et al.
Veröffentlicht: (2025)
von: Ma, Ziyang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Spontaneous Informal Speech Dataset for Punctuation Restoration
von: Liu, Xing Yi, et al.
Veröffentlicht: (2024) -
Carnatic Raga Identification System using Rigorous Time-Delay Neural Network
von: Natesan, Sanjay, et al.
Veröffentlicht: (2024) -
Lla-VAP: LSTM Ensemble of Llama and VAP for Turn-Taking Prediction
von: Jeon, Hyunbae, et al.
Veröffentlicht: (2024) -
Loop Copilot: Conducting AI Ensembles for Music Generation and Iterative Editing
von: Zhang, Yixiao, et al.
Veröffentlicht: (2023) -
NeckCare: Preventing Tech Neck using Hearable-based Multimodal Sensing
von: Chhaglani, Bhawana, et al.
Veröffentlicht: (2024)