Saved in:
Bibliographic Details
Main Authors: Chandra, Shreeram Suresh, Goncalves, Lucas, Lu, Junchen, Busso, Carlos, Sisman, Berrak
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2505.23732
Tags: Add Tag
No Tags, Be the first to tag this record!
Table of Contents:
  • Current emotion-based contrastive language-audio pretraining (CLAP) methods typically learn by naïvely aligning audio samples with corresponding text prompts. Consequently, this approach fails to capture the ordinal nature of emotions, hindering inter-emotion understanding and often resulting in a wide modality gap between the audio and text embeddings due to insufficient alignment. To handle these drawbacks, we introduce EmotionRankCLAP, a supervised contrastive learning approach that uses dimensional attributes of emotional speech and natural language prompts to jointly capture fine-grained emotion variations and improve cross-modal alignment. Our approach utilizes a Rank-N-Contrast objective to learn ordered relationships by contrasting samples based on their rankings in the valence-arousal space. EmotionRankCLAP outperforms existing emotion-CLAP methods in modeling emotion ordinality across modalities, measured via a cross-modal retrieval task.