Foundation Model Embeddings Meet Blended Emotions: A Multimodal Fusion Approach for the BLEMORE Challenge
Fuente:
arXiv
Saved in:
| Main Authors: | Chapariniya, Masoumeh, Farhadipour, Aref, Ebling, Sarah, Dellwo, Volker, Vukovic, Teodora |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Comparative Analysis of Modality Fusion Approaches for Audio-Visual Person Identification and Verification
by: Farhadipour, Aref, et al.
Published: (2024)
by: Farhadipour, Aref, et al.
Published: (2024)
Multimodal Emotion Recognition and Sentiment Analysis in Multi-Party Conversation Contexts
by: Farhadipour, Aref, et al.
Published: (2025)
by: Farhadipour, Aref, et al.
Published: (2025)
Beyond Appearance: Transformer-based Person Identification from Conversational Dynamics
by: Chapariniya, Masoumeh, et al.
Published: (2025)
by: Chapariniya, Masoumeh, et al.
Published: (2025)
Two-Stream Spatial-Temporal Transformer Framework for Person Identification via Natural Conversational Keypoints
by: Chapariniya, Masoumeh, et al.
Published: (2025)
by: Chapariniya, Masoumeh, et al.
Published: (2025)
Micro-Expression-Aware Avatar Fingerprinting via Inter-Frame Feature Differencing
by: Chapariniya, Masoumeh, et al.
Published: (2026)
by: Chapariniya, Masoumeh, et al.
Published: (2026)
Investigating Identity Signals in Conversational Facial Dynamics via Disentangled Expression Features
by: Chapariniya, Masoumeh, et al.
Published: (2025)
by: Chapariniya, Masoumeh, et al.
Published: (2025)
Adaptive Multimodal Person Recognition: A Robust Framework for Handling Missing Modalities
by: Farhadipour, Aref, et al.
Published: (2025)
by: Farhadipour, Aref, et al.
Published: (2025)
Towards Language-Independent Face-Voice Association with Multimodal Foundation Models
by: Farhadipour, Aref, et al.
Published: (2025)
by: Farhadipour, Aref, et al.
Published: (2025)
CL-UZH submission to the NIST SRE 2024 Speaker Recognition Evaluation
by: Farhadipour, Aref, et al.
Published: (2025)
by: Farhadipour, Aref, et al.
Published: (2025)
Not all Blends are Equal: The BLEMORE Dataset of Blended Emotion Expressions with Relative Salience Annotations
by: Lachmann, Tim, et al.
Published: (2026)
by: Lachmann, Tim, et al.
Published: (2026)
Leveraging Self-Supervised Models for Automatic Whispered Speech Recognition
by: Farhadipour, Aref, et al.
Published: (2024)
by: Farhadipour, Aref, et al.
Published: (2024)
EmoLLM: Multimodal Emotional Understanding Meets Large Language Models
by: Yang, Qu, et al.
Published: (2024)
by: Yang, Qu, et al.
Published: (2024)
Deep Neural Networks for Automatic Speaker Recognition Do Not Learn Supra-Segmental Temporal Features
by: Neururer, Daniel, et al.
Published: (2023)
by: Neururer, Daniel, et al.
Published: (2023)
Ordering Matters: Rank-Aware Selective Fusion for Blended Emotion Recognition
by: Lee, Junghyun, et al.
Published: (2026)
by: Lee, Junghyun, et al.
Published: (2026)
Survey of Multimodal Geospatial Foundation Models: Techniques, Applications, and Challenges
by: Yang, Liling, et al.
Published: (2025)
by: Yang, Liling, et al.
Published: (2025)
ChefFusion: Multimodal Foundation Model Integrating Recipe and Food Image Generation
by: Li, Peiyu, et al.
Published: (2024)
by: Li, Peiyu, et al.
Published: (2024)
ECMF: Enhanced Cross-Modal Fusion for Multimodal Emotion Recognition in MER-SEMI Challenge
by: Hu, Juewen, et al.
Published: (2025)
by: Hu, Juewen, et al.
Published: (2025)
BlendFusion -- Scalable Synthetic Data Generation for Diffusion Model Training
by: Venkatesh, Thejas, et al.
Published: (2026)
by: Venkatesh, Thejas, et al.
Published: (2026)
Audio Description Generation in the Era of LLMs and VLMs: A Review of Transferable Generative AI Technologies
by: Gao, Yingqiang, et al.
Published: (2024)
by: Gao, Yingqiang, et al.
Published: (2024)
A Multimodal Fusion Network For Student Emotion Recognition Based on Transformer and Tensor Product
by: Xiang, Ao, et al.
Published: (2024)
by: Xiang, Ao, et al.
Published: (2024)
Sleep Stage Classification using Multimodal Embedding Fusion from EOG and PSM
by: Papillon, Olivier, et al.
Published: (2025)
by: Papillon, Olivier, et al.
Published: (2025)
Anchoring Emotions in Text: Robust Multimodal Fusion for Mimicry Intensity Estimation
by: Zhu, Lingsi, et al.
Published: (2026)
by: Zhu, Lingsi, et al.
Published: (2026)
Expanding the Content-Style Frontier: a Balanced Subspace Blending Approach for Content-Style LoRA Fusion
by: Huang, Linhao
Published: (2025)
by: Huang, Linhao
Published: (2025)
ERIT Lightweight Multimodal Dataset for Elderly Emotion Recognition and Multimodal Fusion Evaluation
by: Frieske, Rita, et al.
Published: (2024)
by: Frieske, Rita, et al.
Published: (2024)
State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition
by: Farhadipour, Aref, et al.
Published: (2025)
by: Farhadipour, Aref, et al.
Published: (2025)
Beyond Imperfections: A Conditional Inpainting Approach for End-to-End Artifact Removal in VTON and Pose Transfer
by: Tabatabaei, Aref, et al.
Published: (2024)
by: Tabatabaei, Aref, et al.
Published: (2024)
Interactive Multimodal Fusion with Temporal Modeling
by: Yu, Jun, et al.
Published: (2025)
by: Yu, Jun, et al.
Published: (2025)
AdaFusion: Prompt-Guided Inference with Adaptive Fusion of Pathology Foundation Models
by: Xiao, Yuxiang, et al.
Published: (2025)
by: Xiao, Yuxiang, et al.
Published: (2025)
HFMF: Hierarchical Fusion Meets Multi-Stream Models for Deepfake Detection
by: Mehta, Anant, et al.
Published: (2025)
by: Mehta, Anant, et al.
Published: (2025)
TeEFusion: Blending Text Embeddings to Distill Classifier-Free Guidance
by: Fu, Minghao, et al.
Published: (2025)
by: Fu, Minghao, et al.
Published: (2025)
TidyVoice 2026 Challenge Evaluation Plan
by: Farhadipour, Aref, et al.
Published: (2026)
by: Farhadipour, Aref, et al.
Published: (2026)
Foundational Question Generation for Video Question Answering via an Embedding-Integrated Approach
by: Oh, Ju-Young
Published: (2025)
by: Oh, Ju-Young
Published: (2025)
Investigating Disability Representations in Text-to-Image Models
by: Tian, Yang, et al.
Published: (2026)
by: Tian, Yang, et al.
Published: (2026)
MANGO: Multimodal Attention-based Normalizing Flow Approach to Fusion Learning
by: Truong, Thanh-Dat, et al.
Published: (2025)
by: Truong, Thanh-Dat, et al.
Published: (2025)
Multimodal Models Meet Presentation Attack Detection on ID Documents
by: Villanueva, Marina, et al.
Published: (2026)
by: Villanueva, Marina, et al.
Published: (2026)
MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled Embedding
by: Liu, Chang, et al.
Published: (2025)
by: Liu, Chang, et al.
Published: (2025)
DreamTalk: When Emotional Talking Head Generation Meets Diffusion Probabilistic Models
by: Ma, Yifeng, et al.
Published: (2023)
by: Ma, Yifeng, et al.
Published: (2023)
Low-Resource Vision Challenges for Foundation Models
by: Zhang, Yunhua, et al.
Published: (2024)
by: Zhang, Yunhua, et al.
Published: (2024)
EmotionHallucer: Evaluating Emotion Hallucinations in Multimodal Large Language Models
by: Xing, Bohao, et al.
Published: (2025)
by: Xing, Bohao, et al.
Published: (2025)
Plug-and-Play Logit Fusion for Heterogeneous Pathology Foundation Models
by: Huang, Gexin, et al.
Published: (2026)
by: Huang, Gexin, et al.
Published: (2026)
Similar Items
-
Comparative Analysis of Modality Fusion Approaches for Audio-Visual Person Identification and Verification
by: Farhadipour, Aref, et al.
Published: (2024) -
Multimodal Emotion Recognition and Sentiment Analysis in Multi-Party Conversation Contexts
by: Farhadipour, Aref, et al.
Published: (2025) -
Beyond Appearance: Transformer-based Person Identification from Conversational Dynamics
by: Chapariniya, Masoumeh, et al.
Published: (2025) -
Two-Stream Spatial-Temporal Transformer Framework for Person Identification via Natural Conversational Keypoints
by: Chapariniya, Masoumeh, et al.
Published: (2025) -
Micro-Expression-Aware Avatar Fingerprinting via Inter-Frame Feature Differencing
by: Chapariniya, Masoumeh, et al.
Published: (2026)