Multimodal Belief Prediction
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Murzaku, John, Soubki, Adil, Rambow, Owen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Synthetic Audio Helps for Cognitive State Tasks
von: Soubki, Adil, et al.
Veröffentlicht: (2025)
von: Soubki, Adil, et al.
Veröffentlicht: (2025)
Analyzing Multimodal Features of Spontaneous Voice Assistant Commands for Mild Cognitive Impairment Detection
von: Lin, Nana, et al.
Veröffentlicht: (2024)
von: Lin, Nana, et al.
Veröffentlicht: (2024)
Towards Robust FastSpeech 2 by Modelling Residual Multimodality
von: Kögel, Fabian, et al.
Veröffentlicht: (2023)
von: Kögel, Fabian, et al.
Veröffentlicht: (2023)
TelME: Teacher-leading Multimodal Fusion Network for Emotion Recognition in Conversation
von: Yun, Taeyang, et al.
Veröffentlicht: (2024)
von: Yun, Taeyang, et al.
Veröffentlicht: (2024)
Towards Early Prediction of Self-Supervised Speech Model Performance
von: Whetten, Ryan, et al.
Veröffentlicht: (2025)
von: Whetten, Ryan, et al.
Veröffentlicht: (2025)
Predicting User Intents and Musical Attributes from Music Discovery Conversations
von: Kwon, Daeyong, et al.
Veröffentlicht: (2024)
von: Kwon, Daeyong, et al.
Veröffentlicht: (2024)
Utilizing Neural Transducers for Two-Stage Text-to-Speech via Semantic Token Prediction
von: Kim, Minchan, et al.
Veröffentlicht: (2024)
von: Kim, Minchan, et al.
Veröffentlicht: (2024)
In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions
von: Fan, Xulin, et al.
Veröffentlicht: (2026)
von: Fan, Xulin, et al.
Veröffentlicht: (2026)
MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models
von: Weck, Benno, et al.
Veröffentlicht: (2024)
von: Weck, Benno, et al.
Veröffentlicht: (2024)
Attentive Fusion: A Transformer-based Approach to Multimodal Hate Speech Detection
von: Mandal, Atanu, et al.
Veröffentlicht: (2024)
von: Mandal, Atanu, et al.
Veröffentlicht: (2024)
ETTA: Elucidating the Design Space of Text-to-Audio Models
von: Lee, Sang-gil, et al.
Veröffentlicht: (2024)
von: Lee, Sang-gil, et al.
Veröffentlicht: (2024)
Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems
von: Park, Taejin, et al.
Veröffentlicht: (2024)
von: Park, Taejin, et al.
Veröffentlicht: (2024)
Introduction to speech recognition
von: Dauphin, Gabriel
Veröffentlicht: (2024)
von: Dauphin, Gabriel
Veröffentlicht: (2024)
AVCap: Leveraging Audio-Visual Features as Text Tokens for Captioning
von: Kim, Jongsuk, et al.
Veröffentlicht: (2024)
von: Kim, Jongsuk, et al.
Veröffentlicht: (2024)
An Analysis of Linear Complexity Attention Substitutes with BEST-RQ
von: Whetten, Ryan, et al.
Veröffentlicht: (2024)
von: Whetten, Ryan, et al.
Veröffentlicht: (2024)
A Closer Look at Neural Codec Resynthesis: Bridging the Gap between Codec and Waveform Generation
von: Liu, Alexander H., et al.
Veröffentlicht: (2024)
von: Liu, Alexander H., et al.
Veröffentlicht: (2024)
On the Semantic Latent Space of Diffusion-Based Text-to-Speech Models
von: Varshavsky-Hassid, Miri, et al.
Veröffentlicht: (2024)
von: Varshavsky-Hassid, Miri, et al.
Veröffentlicht: (2024)
Speech Robust Bench: A Robustness Benchmark For Speech Recognition
von: Shah, Muhammad A., et al.
Veröffentlicht: (2024)
von: Shah, Muhammad A., et al.
Veröffentlicht: (2024)
Property Neurons in Self-Supervised Speech Transformers
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2024)
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2024)
Beyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue Agents
von: Veluri, Bandhav, et al.
Veröffentlicht: (2024)
von: Veluri, Bandhav, et al.
Veröffentlicht: (2024)
Rethinking MUSHRA: Addressing Modern Challenges in Text-to-Speech Evaluation
von: Varadhan, Praveen Srinivasa, et al.
Veröffentlicht: (2024)
von: Varadhan, Praveen Srinivasa, et al.
Veröffentlicht: (2024)
Africa-Centric Self-Supervised Pre-Training for Multilingual Speech Representation in a Sub-Saharan Context
von: Caubrière, Antoine, et al.
Veröffentlicht: (2024)
von: Caubrière, Antoine, et al.
Veröffentlicht: (2024)
Usefulness of Emotional Prosody in Neural Machine Translation
von: Brazier, Charles, et al.
Veröffentlicht: (2024)
von: Brazier, Charles, et al.
Veröffentlicht: (2024)
Conformer-1: Robust ASR via Large-Scale Semisupervised Bootstrapping
von: Zhang, Kevin, et al.
Veröffentlicht: (2024)
von: Zhang, Kevin, et al.
Veröffentlicht: (2024)
XLS-R Deep Learning Model for Multilingual ASR on Low- Resource Languages: Indonesian, Javanese, and Sundanese
von: Arisaputra, Panji, et al.
Veröffentlicht: (2024)
von: Arisaputra, Panji, et al.
Veröffentlicht: (2024)
Exploring Pathological Speech Quality Assessment with ASR-Powered Wav2Vec2 in Data-Scarce Context
von: Nguyen, Tuan, et al.
Veröffentlicht: (2024)
von: Nguyen, Tuan, et al.
Veröffentlicht: (2024)
Energy-Based Models with Applications to Speech and Language Processing
von: Ou, Zhijian
Veröffentlicht: (2024)
von: Ou, Zhijian
Veröffentlicht: (2024)
Disentangling Textual and Acoustic Features of Neural Speech Representations
von: Mohebbi, Hosein, et al.
Veröffentlicht: (2024)
von: Mohebbi, Hosein, et al.
Veröffentlicht: (2024)
Moonshine: Speech Recognition for Live Transcription and Voice Commands
von: Jeffries, Nat, et al.
Veröffentlicht: (2024)
von: Jeffries, Nat, et al.
Veröffentlicht: (2024)
Investigating Disentanglement in a Phoneme-level Speech Codec for Prosody Modeling
von: Karapiperis, Sotirios, et al.
Veröffentlicht: (2024)
von: Karapiperis, Sotirios, et al.
Veröffentlicht: (2024)
Robust and Explainable Depression Identification from Speech Using Vowel-Based Ensemble Learning Approaches
von: Feng, Kexin, et al.
Veröffentlicht: (2024)
von: Feng, Kexin, et al.
Veröffentlicht: (2024)
How Redundant Is the Transformer Stack in Speech Representation Models?
von: Dorszewski, Teresa, et al.
Veröffentlicht: (2024)
von: Dorszewski, Teresa, et al.
Veröffentlicht: (2024)
Papez: Resource-Efficient Speech Separation with Auditory Working Memory
von: Oh, Hyunseok, et al.
Veröffentlicht: (2024)
von: Oh, Hyunseok, et al.
Veröffentlicht: (2024)
A light-weight and efficient punctuation and word casing prediction model for on-device streaming ASR
von: You, Jian, et al.
Veröffentlicht: (2024)
von: You, Jian, et al.
Veröffentlicht: (2024)
Enhancing Out-of-Vocabulary Performance of Indian TTS Systems for Practical Applications through Low-Effort Data Strategies
von: Anand, Srija, et al.
Veröffentlicht: (2024)
von: Anand, Srija, et al.
Veröffentlicht: (2024)
Spontaneous Style Text-to-Speech Synthesis with Controllable Spontaneous Behaviors Based on Language Models
von: Li, Weiqin, et al.
Veröffentlicht: (2024)
von: Li, Weiqin, et al.
Veröffentlicht: (2024)
Meta Learning Text-to-Speech Synthesis in over 7000 Languages
von: Lux, Florian, et al.
Veröffentlicht: (2024)
von: Lux, Florian, et al.
Veröffentlicht: (2024)
CA-SSLR: Condition-Aware Self-Supervised Learning Representation for Generalized Speech Processing
von: Lu, Yen-Ju, et al.
Veröffentlicht: (2024)
von: Lu, Yen-Ju, et al.
Veröffentlicht: (2024)
A multilingual training strategy for low resource Text to Speech
von: Amalas, Asma, et al.
Veröffentlicht: (2024)
von: Amalas, Asma, et al.
Veröffentlicht: (2024)
Understanding Sounds, Missing the Questions: The Challenge of Object Hallucination in Large Audio-Language Models
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2024)
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Synthetic Audio Helps for Cognitive State Tasks
von: Soubki, Adil, et al.
Veröffentlicht: (2025) -
Analyzing Multimodal Features of Spontaneous Voice Assistant Commands for Mild Cognitive Impairment Detection
von: Lin, Nana, et al.
Veröffentlicht: (2024) -
Towards Robust FastSpeech 2 by Modelling Residual Multimodality
von: Kögel, Fabian, et al.
Veröffentlicht: (2023) -
TelME: Teacher-leading Multimodal Fusion Network for Emotion Recognition in Conversation
von: Yun, Taeyang, et al.
Veröffentlicht: (2024) -
Towards Early Prediction of Self-Supervised Speech Model Performance
von: Whetten, Ryan, et al.
Veröffentlicht: (2025)