Quality-Controlled Multimodal Emotion Recognition in Conversations with Identity-Based Transfer Learning and MAMBA Fusion
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Zanxu, Beigi, Homayoon |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Lightweight Self-Supervised Detection of Fundamental Frequency and Accurate Probability of Voicing in Monophonic Music
di: Bitra, Venkat Suprabath, et al.
Pubblicazione: (2026)
di: Bitra, Venkat Suprabath, et al.
Pubblicazione: (2026)
Integrating Spatial and Semantic Embeddings for Stereo Sound Event Localization in Videos
di: Berghi, Davide, et al.
Pubblicazione: (2025)
di: Berghi, Davide, et al.
Pubblicazione: (2025)
Multimodal Biomarkers for Schizophrenia: Towards Individual Symptom Severity Estimation
di: Premananth, Gowtham, et al.
Pubblicazione: (2025)
di: Premananth, Gowtham, et al.
Pubblicazione: (2025)
Fast Swap-Based Element Selection for Multiplication-Free Dimension Reduction
di: Ono, Nobutaka
Pubblicazione: (2026)
di: Ono, Nobutaka
Pubblicazione: (2026)
Leveraging Reverberation and Visual Depth Cues for Sound Event Localization and Detection with Distance Estimation
di: Berghi, Davide, et al.
Pubblicazione: (2024)
di: Berghi, Davide, et al.
Pubblicazione: (2024)
W4S4: WaLRUS Meets S4 for Long-Range Sequence Modeling
di: Babaei, Hossein, et al.
Pubblicazione: (2025)
di: Babaei, Hossein, et al.
Pubblicazione: (2025)
Spatial and Semantic Embedding Integration for Stereo Sound Event Localization and Detection in Regular Videos
di: Berghi, Davide, et al.
Pubblicazione: (2025)
di: Berghi, Davide, et al.
Pubblicazione: (2025)
SaFARi: State-Space Models for Frame-Agnostic Representation
di: Babaei, Hossein, et al.
Pubblicazione: (2025)
di: Babaei, Hossein, et al.
Pubblicazione: (2025)
Leveraging Unlabeled Audio-Visual Data in Speech Emotion Recognition using Knowledge Distillation
di: Pendyala, Varsha, et al.
Pubblicazione: (2025)
di: Pendyala, Varsha, et al.
Pubblicazione: (2025)
Multimodal sensor fusion for real-time location-dependent defect detection in laser-directed energy deposition
di: Chen, Lequn, et al.
Pubblicazione: (2023)
di: Chen, Lequn, et al.
Pubblicazione: (2023)
ToS: A Team of Specialists ensemble framework for Stereo Sound Event Localization and Detection with distance estimation in Video
di: Berghi, Davide, et al.
Pubblicazione: (2026)
di: Berghi, Davide, et al.
Pubblicazione: (2026)
Livestock feeding behaviour: A review on automated systems for ruminant monitoring
di: Chelotti, José, et al.
Pubblicazione: (2023)
di: Chelotti, José, et al.
Pubblicazione: (2023)
LiVeAction: a Lightweight, Versatile, and Asymmetric Neural Codec Design for Real-time Operation
di: Jacobellis, Dan, et al.
Pubblicazione: (2026)
di: Jacobellis, Dan, et al.
Pubblicazione: (2026)
Video Soundtrack Generation by Aligning Emotions and Temporal Boundaries
di: Sulun, Serkan, et al.
Pubblicazione: (2025)
di: Sulun, Serkan, et al.
Pubblicazione: (2025)
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet
di: Zhong, Zhi, et al.
Pubblicazione: (2025)
di: Zhong, Zhi, et al.
Pubblicazione: (2025)
A Smart-Glasses for Emergency Medical Services via Multimodal Multitask Learning
di: Jin, Liuyi, et al.
Pubblicazione: (2025)
di: Jin, Liuyi, et al.
Pubblicazione: (2025)
WaLRUS: Wavelets for Long-range Representation Using SSMs
di: Babaei, Hossein, et al.
Pubblicazione: (2025)
di: Babaei, Hossein, et al.
Pubblicazione: (2025)
SIM: Surface-based fMRI Analysis for Inter-Subject Multimodal Decoding from Movie-Watching Experiments
di: Dahan, Simon, et al.
Pubblicazione: (2025)
di: Dahan, Simon, et al.
Pubblicazione: (2025)
Carnatic Raga Identification System using Rigorous Time-Delay Neural Network
di: Natesan, Sanjay, et al.
Pubblicazione: (2024)
di: Natesan, Sanjay, et al.
Pubblicazione: (2024)
Learned Compression for Compressed Learning
di: Jacobellis, Dan, et al.
Pubblicazione: (2024)
di: Jacobellis, Dan, et al.
Pubblicazione: (2024)
Multimodal Marvels of Deep Learning in Medical Diagnosis: A Comprehensive Review of COVID-19 Detection
di: Islam, Md Shofiqul, et al.
Pubblicazione: (2025)
di: Islam, Md Shofiqul, et al.
Pubblicazione: (2025)
RestoreGrad: Signal Restoration Using Conditional Denoising Diffusion Models with Jointly Learned Prior
di: Lee, Ching-Hua, et al.
Pubblicazione: (2025)
di: Lee, Ching-Hua, et al.
Pubblicazione: (2025)
Spoken question answering for visual queries
di: Shabtay, Nimrod, et al.
Pubblicazione: (2025)
di: Shabtay, Nimrod, et al.
Pubblicazione: (2025)
Efficient Test-Time Adaptation through Latent Subspace Coefficients Search
di: Luo, Xinyu, et al.
Pubblicazione: (2025)
di: Luo, Xinyu, et al.
Pubblicazione: (2025)
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization
di: Malard, Hugo, et al.
Pubblicazione: (2024)
di: Malard, Hugo, et al.
Pubblicazione: (2024)
SoundSil-DS: Deep Denoising and Segmentation of Sound-field Images with Silhouettes
di: Tanigawa, Risako, et al.
Pubblicazione: (2024)
di: Tanigawa, Risako, et al.
Pubblicazione: (2024)
A multi-modal approach for identifying schizophrenia using cross-modal attention
di: Premananth, Gowtham, et al.
Pubblicazione: (2023)
di: Premananth, Gowtham, et al.
Pubblicazione: (2023)
Learning Video Temporal Dynamics with Cross-Modal Attention for Robust Audio-Visual Speech Recognition
di: Kim, Sungnyun, et al.
Pubblicazione: (2024)
di: Kim, Sungnyun, et al.
Pubblicazione: (2024)
Attentive AV-FusionNet: Audio-Visual Quality Prediction with Hybrid Attention
di: Salaj, Ina, et al.
Pubblicazione: (2025)
di: Salaj, Ina, et al.
Pubblicazione: (2025)
End-to-end audio-visual learning for cochlear implant sound coding simulations in noisy environments
di: Lin, Meng-Ping, et al.
Pubblicazione: (2025)
di: Lin, Meng-Ping, et al.
Pubblicazione: (2025)
Audio-Visual Approach For Multimodal Concurrent Speaker Detection
di: Eliav, Amit, et al.
Pubblicazione: (2024)
di: Eliav, Amit, et al.
Pubblicazione: (2024)
Multimodal Machine Learning Can Predict Videoconference Fluidity and Enjoyment
di: Chang, Andrew, et al.
Pubblicazione: (2025)
di: Chang, Andrew, et al.
Pubblicazione: (2025)
Blind Separation of Vibration Sources using Deep Learning and Deconvolution
di: Makienko, Igor, et al.
Pubblicazione: (2024)
di: Makienko, Igor, et al.
Pubblicazione: (2024)
Deep Learning for Steganalysis of Diverse Data Types: A review of methods, taxonomy, challenges and future directions
di: Kheddar, Hamza, et al.
Pubblicazione: (2023)
di: Kheddar, Hamza, et al.
Pubblicazione: (2023)
Reacting like Humans: Incorporating Intrinsic Human Behaviors into NAO through Sound-Based Reactions to Fearful and Shocking Events for Enhanced Sociability
di: Ghadami, Ali, et al.
Pubblicazione: (2023)
di: Ghadami, Ali, et al.
Pubblicazione: (2023)
Scaling Transformers for Low-Bitrate High-Quality Speech Coding
di: Parker, Julian D, et al.
Pubblicazione: (2024)
di: Parker, Julian D, et al.
Pubblicazione: (2024)
Laugh Now Cry Later: Controlling Time-Varying Emotional States of Flow-Matching-Based Zero-Shot Text-to-Speech
di: Wu, Haibin, et al.
Pubblicazione: (2024)
di: Wu, Haibin, et al.
Pubblicazione: (2024)
When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds
di: Kang, Minsu, et al.
Pubblicazione: (2025)
di: Kang, Minsu, et al.
Pubblicazione: (2025)
Language Bias in Self-Supervised Learning For Automatic Speech Recognition
di: Storey, Edward, et al.
Pubblicazione: (2025)
di: Storey, Edward, et al.
Pubblicazione: (2025)
TinyChirp: Bird Song Recognition Using TinyML Models on Low-power Wireless Acoustic Sensors
di: Huang, Zhaolan, et al.
Pubblicazione: (2024)
di: Huang, Zhaolan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Lightweight Self-Supervised Detection of Fundamental Frequency and Accurate Probability of Voicing in Monophonic Music
di: Bitra, Venkat Suprabath, et al.
Pubblicazione: (2026) -
Integrating Spatial and Semantic Embeddings for Stereo Sound Event Localization in Videos
di: Berghi, Davide, et al.
Pubblicazione: (2025) -
Multimodal Biomarkers for Schizophrenia: Towards Individual Symptom Severity Estimation
di: Premananth, Gowtham, et al.
Pubblicazione: (2025) -
Fast Swap-Based Element Selection for Multiplication-Free Dimension Reduction
di: Ono, Nobutaka
Pubblicazione: (2026) -
Leveraging Reverberation and Visual Depth Cues for Sound Event Localization and Detection with Distance Estimation
di: Berghi, Davide, et al.
Pubblicazione: (2024)