MoralCLIP: Contrastive Alignment of Vision-and-Language Representations with Moral Foundations Theory

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Condez, Ana Carolina, Tavares, Diogo, Magalhães, João
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915585889140736
author Condez, Ana Carolina
Tavares, Diogo
Magalhães, João
author_facet Condez, Ana Carolina
Tavares, Diogo
Magalhães, João
contents Recent advances in vision-language models have enabled rich semantic understanding across modalities. However, these encoding methods lack the ability to interpret or reason about the moral dimensions of content-a crucial aspect of human cognition. In this paper, we address this gap by introducing MoralCLIP, a novel embedding representation method that extends multimodal learning with explicit moral grounding based on Moral Foundations Theory (MFT). Our approach integrates visual and textual moral cues into a unified embedding space, enabling cross-modal moral alignment. MoralCLIP is grounded on the multi-label dataset Social-Moral Image Database to identify co-occurring moral foundations in visual content. For MoralCLIP training, we design a moral data augmentation strategy to scale our annotated dataset to 15,000 image-text pairs labeled with MFT-aligned dimensions. Our results demonstrate that explicit moral supervision improves both unimodal and multimodal understanding of moral content, establishing a foundation for morally-aware AI systems capable of recognizing and aligning with human moral values.
format Preprint
id arxiv_https___arxiv_org_abs_2506_05696
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MoralCLIP: Contrastive Alignment of Vision-and-Language Representations with Moral Foundations Theory
Condez, Ana Carolina
Tavares, Diogo
Magalhães, João
Computer Vision and Pattern Recognition
Recent advances in vision-language models have enabled rich semantic understanding across modalities. However, these encoding methods lack the ability to interpret or reason about the moral dimensions of content-a crucial aspect of human cognition. In this paper, we address this gap by introducing MoralCLIP, a novel embedding representation method that extends multimodal learning with explicit moral grounding based on Moral Foundations Theory (MFT). Our approach integrates visual and textual moral cues into a unified embedding space, enabling cross-modal moral alignment. MoralCLIP is grounded on the multi-label dataset Social-Moral Image Database to identify co-occurring moral foundations in visual content. For MoralCLIP training, we design a moral data augmentation strategy to scale our annotated dataset to 15,000 image-text pairs labeled with MFT-aligned dimensions. Our results demonstrate that explicit moral supervision improves both unimodal and multimodal understanding of moral content, establishing a foundation for morally-aware AI systems capable of recognizing and aligning with human moral values.
title MoralCLIP: Contrastive Alignment of Vision-and-Language Representations with Moral Foundations Theory
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.05696