MM-SCALE: Grounded Multimodal Moral Reasoning via Scalar Judgment and Listwise Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Park, Eunkyu, Deng, Wesley Hanwen, Jin, Cheyon, Maldaner, Matheus Kunzler, Wheeler, Jordan, Hong, Jason I., Shen, Hong, Perer, Adam, Holstein, Ken, Eslami, Motahhare, Kim, Gunhee
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910010499399680
author Park, Eunkyu
Deng, Wesley Hanwen
Jin, Cheyon
Maldaner, Matheus Kunzler
Wheeler, Jordan
Hong, Jason I.
Shen, Hong
Perer, Adam
Holstein, Ken
Eslami, Motahhare
Kim, Gunhee
author_facet Park, Eunkyu
Deng, Wesley Hanwen
Jin, Cheyon
Maldaner, Matheus Kunzler
Wheeler, Jordan
Hong, Jason I.
Shen, Hong
Perer, Adam
Holstein, Ken
Eslami, Motahhare
Kim, Gunhee
contents Vision-Language Models (VLMs) continue to struggle to make morally salient judgments in multimodal and socially ambiguous contexts. Prior works typically rely on binary or pairwise supervision, which often fail to capture the continuous and pluralistic nature of human moral reasoning. We present MM-SCALE (Multimodal Moral Scale), a large-scale dataset for aligning VLMs with human moral preferences through 5-point scalar ratings and explicit modality grounding. Each image-scenario pair is annotated with moral acceptability scores and grounded reasoning labels by humans using an interface we tailored for data collection, enabling listwise preference optimization over ranked scenario sets. By moving from discrete to scalar supervision, our framework provides richer alignment signals and finer calibration of multimodal moral reasoning. Experiments show that VLMs fine-tuned on MM-SCALE achieve higher ranking fidelity and more stable safety calibration than those trained with binary signals.
format Preprint
id arxiv_https___arxiv_org_abs_2602_03665
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MM-SCALE: Grounded Multimodal Moral Reasoning via Scalar Judgment and Listwise Alignment
Park, Eunkyu
Deng, Wesley Hanwen
Jin, Cheyon
Maldaner, Matheus Kunzler
Wheeler, Jordan
Hong, Jason I.
Shen, Hong
Perer, Adam
Holstein, Ken
Eslami, Motahhare
Kim, Gunhee
Computer Vision and Pattern Recognition
Human-Computer Interaction
Vision-Language Models (VLMs) continue to struggle to make morally salient judgments in multimodal and socially ambiguous contexts. Prior works typically rely on binary or pairwise supervision, which often fail to capture the continuous and pluralistic nature of human moral reasoning. We present MM-SCALE (Multimodal Moral Scale), a large-scale dataset for aligning VLMs with human moral preferences through 5-point scalar ratings and explicit modality grounding. Each image-scenario pair is annotated with moral acceptability scores and grounded reasoning labels by humans using an interface we tailored for data collection, enabling listwise preference optimization over ranked scenario sets. By moving from discrete to scalar supervision, our framework provides richer alignment signals and finer calibration of multimodal moral reasoning. Experiments show that VLMs fine-tuned on MM-SCALE achieve higher ranking fidelity and more stable safety calibration than those trained with binary signals.
title MM-SCALE: Grounded Multimodal Moral Reasoning via Scalar Judgment and Listwise Alignment
topic Computer Vision and Pattern Recognition
Human-Computer Interaction
url https://arxiv.org/abs/2602.03665