MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions
Fuente:
arXiv
Saved in:
| Main Authors: | Chatzichristodoulou, Georgios, Kosmopoulou, Despoina, Kritikos, Antonios, Poulopoulou, Anastasia, Georgiou, Efthymios, Katsamanis, Athanasios, Katsouros, Vassilis, Potamianos, Alexandros |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Masked Diffusion Language Models with Frequency-Informed Training
by: Kosmopoulou, Despoina, et al.
Published: (2025)
by: Kosmopoulou, Despoina, et al.
Published: (2025)
DeepMLF: Multimodal language model with learnable tokens for deep fusion in sentiment analysis
by: Georgiou, Efthymios, et al.
Published: (2025)
by: Georgiou, Efthymios, et al.
Published: (2025)
The Greek podcast corpus: Competitive speech models for low-resourced languages with weakly supervised data
by: Paraskevopoulos, Georgios, et al.
Published: (2024)
by: Paraskevopoulos, Georgios, et al.
Published: (2024)
Y-Drop: A Conductance based Dropout for fully connected layers
by: Georgiou, Efthymios, et al.
Published: (2024)
by: Georgiou, Efthymios, et al.
Published: (2024)
VOX-KRIKRI: Unifying Speech and Language through Continuous Fusion
by: Damianos, Dimitrios, et al.
Published: (2025)
by: Damianos, Dimitrios, et al.
Published: (2025)
Krikri: Advancing Open Large Language Models for Greek
by: Roussis, Dimitris, et al.
Published: (2025)
by: Roussis, Dimitris, et al.
Published: (2025)
Meltemi: The first open Large Language Model for Greek
by: Voukoutis, Leon, et al.
Published: (2024)
by: Voukoutis, Leon, et al.
Published: (2024)
The Strong Pull of Prior Knowledge in Large Language Models and Its Impact on Emotion Recognition
by: Chochlakis, Georgios, et al.
Published: (2024)
by: Chochlakis, Georgios, et al.
Published: (2024)
Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025
by: Ferreira, Alef Iury Siqueira, et al.
Published: (2025)
by: Ferreira, Alef Iury Siqueira, et al.
Published: (2025)
CrossFlowDG: Bridging the Modality Gap with Cross-modal Flow Matching for Domain Generalization
by: Kritikos, Antonios, et al.
Published: (2026)
by: Kritikos, Antonios, et al.
Published: (2026)
ABHINAYA -- A System for Speech Emotion Recognition In Naturalistic Conditions Challenge
by: Dutta, Soumya, et al.
Published: (2025)
by: Dutta, Soumya, et al.
Published: (2025)
EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing
by: Sioros, Vassilis, et al.
Published: (2025)
by: Sioros, Vassilis, et al.
Published: (2025)
Spira: Exploiting Voxel Data Structural Properties for Efficient Sparse Convolution in Point Cloud Networks
by: Adamopoulos, Dionysios, et al.
Published: (2025)
by: Adamopoulos, Dionysios, et al.
Published: (2025)
Auto-Compressing Networks
by: Dorovatas, Vaggelis, et al.
Published: (2025)
by: Dorovatas, Vaggelis, et al.
Published: (2025)
MSDA: Combining Pseudo-labeling and Self-Supervision for Unsupervised Domain Adaptation in ASR
by: Damianos, Dimitrios, et al.
Published: (2025)
by: Damianos, Dimitrios, et al.
Published: (2025)
Developing a High-performance Framework for Speech Emotion Recognition in Naturalistic Conditions Challenge for Emotional Attribute Prediction
by: Lertpetchpun, Thanathai, et al.
Published: (2025)
by: Lertpetchpun, Thanathai, et al.
Published: (2025)
Transcoders Trace Visual Grounding and Hallucinations in Vision-Language Models
by: Damianos, Dimitrios, et al.
Published: (2026)
by: Damianos, Dimitrios, et al.
Published: (2026)
BloomWise: Enhancing Problem-Solving capabilities of Large Language Models using Bloom's-Taxonomy-Inspired Prompts
by: Zoumpoulidi, Maria-Eleni, et al.
Published: (2024)
by: Zoumpoulidi, Maria-Eleni, et al.
Published: (2024)
Aggregation Artifacts in Subjective Tasks Collapse Large Language Models' Posteriors
by: Chochlakis, Georgios, et al.
Published: (2024)
by: Chochlakis, Georgios, et al.
Published: (2024)
Age-Inclusive 3D Human Mesh Recovery for Action-Preserving Data Anonymization
by: Chatzichristodoulou, Georgios, et al.
Published: (2025)
by: Chatzichristodoulou, Georgios, et al.
Published: (2025)
Analytical solution of the Poiseuille flow of a De Kee viscoplastic fluid
by: Syrakos, Alexandros, et al.
Published: (2024)
by: Syrakos, Alexandros, et al.
Published: (2024)
A revisit of the development of viscoplastic flow in pipes and channels
by: Syrakos, Alexandros, et al.
Published: (2024)
by: Syrakos, Alexandros, et al.
Published: (2024)
Student‐Created Digital Stories in Primary School, Interpreting the Day/Night Cycle Through Conceptual Representations
by: Georgios Kritikos, et al.
Published: (2025)
by: Georgios Kritikos, et al.
Published: (2025)
Exposing Hidden Biases in Text-to-Image Models via Automated Prompt Search
by: Plitsis, Manos, et al.
Published: (2025)
by: Plitsis, Manos, et al.
Published: (2025)
LC-Protonets: Multi-Label Few-Shot Learning for World Music Audio Tagging
by: Papaioannou, Charilaos, et al.
Published: (2024)
by: Papaioannou, Charilaos, et al.
Published: (2024)
Universal Music Representations? Evaluating Foundation Models on World Music Corpora
by: Papaioannou, Charilaos, et al.
Published: (2025)
by: Papaioannou, Charilaos, et al.
Published: (2025)
Bimodal Connection Attention Fusion for Speech Emotion Recognition
by: Luo, Jiachen, et al.
Published: (2025)
by: Luo, Jiachen, et al.
Published: (2025)
A Transformer-Based Framework for Greek Sign Language Production using Extended Skeletal Motion Representations
by: Pratikaki, Chrysa, et al.
Published: (2025)
by: Pratikaki, Chrysa, et al.
Published: (2025)
WavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion Recognition
by: Li, Feng, et al.
Published: (2024)
by: Li, Feng, et al.
Published: (2024)
MSP-Conversation: A Corpus for Naturalistic, Time-Continuous Emotion Recognition
by: Martinez-Lucas, Luz, et al.
Published: (2026)
by: Martinez-Lucas, Luz, et al.
Published: (2026)
Semantic F1 Scores: Fair Evaluation Under Fuzzy Class Boundaries
by: Chochlakis, Georgios, et al.
Published: (2025)
by: Chochlakis, Georgios, et al.
Published: (2025)
CultureMERT: Continual Pre-Training for Cross-Cultural Music Representation Learning
by: Kanatas, Angelos-Nikolaos, et al.
Published: (2025)
by: Kanatas, Angelos-Nikolaos, et al.
Published: (2025)
Probing Multimodal Fusion in the Brain: The Dominance of Audiovisual Streams in Naturalistic Encoding
by: Abdollahi, Hamid, et al.
Published: (2025)
by: Abdollahi, Hamid, et al.
Published: (2025)
Explainable Transformer-CNN Fusion for Noise-Robust Speech Emotion Recognition
by: Chakrabarty, Sudip, et al.
Published: (2025)
by: Chakrabarty, Sudip, et al.
Published: (2025)
Spaces of multiscaled lines with collision
by: Robotis, Antonios-Alexandros
Published: (2024)
by: Robotis, Antonios-Alexandros
Published: (2024)
Admissible subcategories of noncommutative curves
by: Robotis, Antonios-Alexandros
Published: (2023)
by: Robotis, Antonios-Alexandros
Published: (2023)
ChatGPT produces more "lazy" thinkers: Evidence of cognitive engagement decline
by: Georgiou, Georgios P.
Published: (2025)
by: Georgiou, Georgios P.
Published: (2025)
Capabilities of GPT-5 across critical domains: Is it the next breakthrough?
by: Georgiou, Georgios P.
Published: (2025)
by: Georgiou, Georgios P.
Published: (2025)
Differentiating Between Human-Written and AI-Generated Texts Using Automatically Extracted Linguistic Features
by: Georgiou, Georgios P.
Published: (2024)
by: Georgiou, Georgios P.
Published: (2024)
Enhancing nonnative speech perception and production through an AI-powered application
by: Georgiou, Georgios P.
Published: (2025)
by: Georgiou, Georgios P.
Published: (2025)
Similar Items
-
Masked Diffusion Language Models with Frequency-Informed Training
by: Kosmopoulou, Despoina, et al.
Published: (2025) -
DeepMLF: Multimodal language model with learnable tokens for deep fusion in sentiment analysis
by: Georgiou, Efthymios, et al.
Published: (2025) -
The Greek podcast corpus: Competitive speech models for low-resourced languages with weakly supervised data
by: Paraskevopoulos, Georgios, et al.
Published: (2024) -
Y-Drop: A Conductance based Dropout for fully connected layers
by: Georgiou, Efthymios, et al.
Published: (2024) -
VOX-KRIKRI: Unifying Speech and Language through Continuous Fusion
by: Damianos, Dimitrios, et al.
Published: (2025)