Multimodal Contextualized Semantic Parsing from Speech
Fuente:
arXiv
Saved in:
| Main Authors: | Voas, Jordan, Mooney, Raymond, Harwath, David |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Multimodal Emotion Recognition System: Integrating Facial Expressions, Body Movement, Speech, and Spoken Language
by: Kraack, Kris
Published: (2024)
by: Kraack, Kris
Published: (2024)
Measuring Sound Symbolism in Audio-visual Models
by: Tseng, Wei-Cheng, et al.
Published: (2024)
by: Tseng, Wei-Cheng, et al.
Published: (2024)
Clip-TTS: Contrastive Text-content and Mel-spectrogram, A High-Quality Text-to-Speech Method based on Contextual Semantic Understanding
by: Liu, Tianyun
Published: (2025)
by: Liu, Tianyun
Published: (2025)
Spontaneous Informal Speech Dataset for Punctuation Restoration
by: Liu, Xing Yi, et al.
Published: (2024)
by: Liu, Xing Yi, et al.
Published: (2024)
Abjad-Kids: An Arabic Speech Classification Dataset for Primary Education
by: Snoubara, Abdul Aziz, et al.
Published: (2026)
by: Snoubara, Abdul Aziz, et al.
Published: (2026)
VoXtream: Full-Stream Text-to-Speech with Extremely Low Latency
by: Torgashov, Nikita, et al.
Published: (2025)
by: Torgashov, Nikita, et al.
Published: (2025)
AIx Speed: Playback Speed Optimization Using Listening Comprehension of Speech Recognition Models
by: Kawamura, Kazuki, et al.
Published: (2024)
by: Kawamura, Kazuki, et al.
Published: (2024)
Deep Learning Models in Speech Recognition: Measuring GPU Energy Consumption, Impact of Noise and Model Quantization for Edge Deployment
by: Chakravarty, Aditya
Published: (2024)
by: Chakravarty, Aditya
Published: (2024)
Efficient Ensemble for Multimodal Punctuation Restoration using Time-Delay Neural Network
by: Liu, Xing Yi, et al.
Published: (2023)
by: Liu, Xing Yi, et al.
Published: (2023)
HRTF upsampling with a generative adversarial network using a gnomonic equiangular projection
by: Hogg, Aidan O. T., et al.
Published: (2023)
by: Hogg, Aidan O. T., et al.
Published: (2023)
A Joint Cross-Attention Model for Audio-Visual Fusion in Dimensional Emotion Recognition
by: Praveen, R. Gnana, et al.
Published: (2022)
by: Praveen, R. Gnana, et al.
Published: (2022)
Speech vs. Transcript: Does It Matter for Human Annotators in Speech Summarization?
by: Sharma, Roshan, et al.
Published: (2024)
by: Sharma, Roshan, et al.
Published: (2024)
Detecting COPD Through Speech Analysis: A Dataset of Danish Speech and Machine Learning Approach
by: Sankey-Olsen, Cuno, et al.
Published: (2025)
by: Sankey-Olsen, Cuno, et al.
Published: (2025)
Artificial Neural Networks to Recognize Speakers Division from Continuous Bengali Speech
by: Ali, Hasmot, et al.
Published: (2024)
by: Ali, Hasmot, et al.
Published: (2024)
Optimizing Multilingual Text-To-Speech with Accents & Emotions
by: Pawar, Pranav, et al.
Published: (2025)
by: Pawar, Pranav, et al.
Published: (2025)
Human Feedback Driven Dynamic Speech Emotion Recognition
by: Fedorov, Ilya, et al.
Published: (2025)
by: Fedorov, Ilya, et al.
Published: (2025)
Loop Copilot: Conducting AI Ensembles for Music Generation and Iterative Editing
by: Zhang, Yixiao, et al.
Published: (2023)
by: Zhang, Yixiao, et al.
Published: (2023)
A conversational gesture synthesis system based on emotions and semantics
by: Hoang-Minh, Thanh
Published: (2025)
by: Hoang-Minh, Thanh
Published: (2025)
LLAMAPIE: Proactive In-Ear Conversation Assistants
by: Chen, Tuochao, et al.
Published: (2025)
by: Chen, Tuochao, et al.
Published: (2025)
VoXtream2: Full-stream TTS with dynamic speaking rate control
by: Torgashov, Nikita, et al.
Published: (2026)
by: Torgashov, Nikita, et al.
Published: (2026)
Literary and Colloquial Tamil Dialect Identification
by: Nanmalar, M., et al.
Published: (2024)
by: Nanmalar, M., et al.
Published: (2024)
EgoTrigger: Toward Audio-Driven Image Capture for Human Memory Enhancement in All-Day Energy-Efficient Smart Glasses
by: Paruchuri, Akshay, et al.
Published: (2025)
by: Paruchuri, Akshay, et al.
Published: (2025)
A Cloud-Based Cross-Modal Transformer for Emotion Recognition and Adaptive Human-Computer Interaction
by: Zhong, Ziwen, et al.
Published: (2025)
by: Zhong, Ziwen, et al.
Published: (2025)
Soundify: Matching Sound Effects to Video
by: Lin, David Chuan-En, et al.
Published: (2021)
by: Lin, David Chuan-En, et al.
Published: (2021)
InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-training
by: Wang, Dingdong, et al.
Published: (2025)
by: Wang, Dingdong, et al.
Published: (2025)
SonicSieve: Bringing Directional Speech Extraction to Smartphones Using Acoustic Microstructures
by: Yuan, Kuang, et al.
Published: (2025)
by: Yuan, Kuang, et al.
Published: (2025)
Investigating the Effects of Large-Scale Pseudo-Stereo Data and Different Speech Foundation Model on Dialogue Generative Spoken Language Model
by: Fu, Yu-Kuan, et al.
Published: (2024)
by: Fu, Yu-Kuan, et al.
Published: (2024)
HiCMAE: Hierarchical Contrastive Masked Autoencoder for Self-Supervised Audio-Visual Emotion Recognition
by: Sun, Licai, et al.
Published: (2024)
by: Sun, Licai, et al.
Published: (2024)
Integrating Representational Gestures into Automatically Generated Embodied Explanations and its Effects on Understanding and Interaction Quality
by: Robrecht, Amelie Sophie, et al.
Published: (2024)
by: Robrecht, Amelie Sophie, et al.
Published: (2024)
Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese
by: Wang, Xihuai, et al.
Published: (2025)
by: Wang, Xihuai, et al.
Published: (2025)
Multimodal Segmentation for Vocal Tract Modeling
by: Jain, Rishi, et al.
Published: (2024)
by: Jain, Rishi, et al.
Published: (2024)
Speech-driven Personalized Gesture Synthetics: Harnessing Automatic Fuzzy Feature Inference
by: Zhang, Fan, et al.
Published: (2024)
by: Zhang, Fan, et al.
Published: (2024)
JEP-KD: Joint-Embedding Predictive Architecture Based Knowledge Distillation for Visual Speech Recognition
by: Sun, Chang, et al.
Published: (2024)
by: Sun, Chang, et al.
Published: (2024)
BanglaRobustNet: A Hybrid Denoising-Attention Architecture for Robust Bangla Speech Recognition
by: Ridoy, Md Sazzadul Islam, et al.
Published: (2026)
by: Ridoy, Md Sazzadul Islam, et al.
Published: (2026)
2D or not 2D: How Does the Dimensionality of Gesture Representation Affect 3D Co-Speech Gesture Generation?
by: Guichoux, Téo, et al.
Published: (2024)
by: Guichoux, Téo, et al.
Published: (2024)
Detecting the terminality of speech-turn boundary for spoken interactions in French TV and Radio content
by: Uro, Rémi, et al.
Published: (2024)
by: Uro, Rémi, et al.
Published: (2024)
[b]=[d]-[t]+[p]: Self-supervised Speech Models Discover Phonological Vector Arithmetic
by: Choi, Kwanghee, et al.
Published: (2026)
by: Choi, Kwanghee, et al.
Published: (2026)
MUSE: A Run-Centric Platform for Multimodal Unified Safety Evaluation of Large Language Models
by: Wang, Zhongxi, et al.
Published: (2026)
by: Wang, Zhongxi, et al.
Published: (2026)
Voice Passing : a Non-Binary Voice Gender Prediction System for evaluating Transgender voice transition
by: Doukhan, David, et al.
Published: (2024)
by: Doukhan, David, et al.
Published: (2024)
Self-Supervised Speech Models Encode Phonetic Context via Position-dependent Orthogonal Subspaces
by: Choi, Kwanghee, et al.
Published: (2026)
by: Choi, Kwanghee, et al.
Published: (2026)
Similar Items
-
A Multimodal Emotion Recognition System: Integrating Facial Expressions, Body Movement, Speech, and Spoken Language
by: Kraack, Kris
Published: (2024) -
Measuring Sound Symbolism in Audio-visual Models
by: Tseng, Wei-Cheng, et al.
Published: (2024) -
Clip-TTS: Contrastive Text-content and Mel-spectrogram, A High-Quality Text-to-Speech Method based on Contextual Semantic Understanding
by: Liu, Tianyun
Published: (2025) -
Spontaneous Informal Speech Dataset for Punctuation Restoration
by: Liu, Xing Yi, et al.
Published: (2024) -
Abjad-Kids: An Arabic Speech Classification Dataset for Primary Education
by: Snoubara, Abdul Aziz, et al.
Published: (2026)