Do VLMs Have a Moral Backbone? A Study on the Fragile Morality of Vision-Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Zhining, Wang, Tianyi, Lin, Xiao, Ouyang, Penghao, Li, Gaotang, Yang, Ze, Liu, Hui, Keswani, Sumit, Pardeshi, Vishwa, Zhao, Huijun, Fan, Wei, Tong, Hanghang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MORALISE: A Structured Benchmark for Moral Alignment in Visual Language Models
di: Lin, Xiao, et al.
Pubblicazione: (2025)
di: Lin, Xiao, et al.
Pubblicazione: (2025)
Taming Knowledge Conflicts in Language Models
di: Li, Gaotang, et al.
Pubblicazione: (2025)
di: Li, Gaotang, et al.
Pubblicazione: (2025)
LegalDrill: Diagnosis-Driven Synthesis for Legal Reasoning in Small Language Models
di: Li, Tianchun, et al.
Pubblicazione: (2026)
di: Li, Tianchun, et al.
Pubblicazione: (2026)
On the Pros and Cons of Active Learning for Moral Preference Elicitation
di: Keswani, Vijay, et al.
Pubblicazione: (2024)
di: Keswani, Vijay, et al.
Pubblicazione: (2024)
Seeing but Not Believing: Probing the Disconnect Between Visual Attention and Answer Correctness in VLMs
di: Liu, Zhining, et al.
Pubblicazione: (2025)
di: Liu, Zhining, et al.
Pubblicazione: (2025)
On The Stability of Moral Preferences: A Problem with Computational Elicitation Methods
di: Boerstler, Kyle, et al.
Pubblicazione: (2024)
di: Boerstler, Kyle, et al.
Pubblicazione: (2024)
The Fragility Of Moral Judgment In Large Language Models
di: van Nuenen, Tom, et al.
Pubblicazione: (2026)
di: van Nuenen, Tom, et al.
Pubblicazione: (2026)
SelfElicit: Your Language Model Secretly Knows Where is the Relevant Evidence
di: Liu, Zhining, et al.
Pubblicazione: (2025)
di: Liu, Zhining, et al.
Pubblicazione: (2025)
Moral Change or Noise? On Problems of Aligning AI With Temporally Unstable Human Feedback
di: Keswani, Vijay, et al.
Pubblicazione: (2025)
di: Keswani, Vijay, et al.
Pubblicazione: (2025)
Have the VLMs Lost Confidence? A Study of Sycophancy in VLMs
di: Li, Shuo, et al.
Pubblicazione: (2024)
di: Li, Shuo, et al.
Pubblicazione: (2024)
Can AI Model the Complexities of Human Moral Decision-Making? A Qualitative Study of Kidney Allocation Decisions
di: Keswani, Vijay, et al.
Pubblicazione: (2025)
di: Keswani, Vijay, et al.
Pubblicazione: (2025)
"Pull or Not to Pull?'': Investigating Moral Biases in Leading Large Language Models Across Ethical Dilemmas
di: Ding, Junchen, et al.
Pubblicazione: (2025)
di: Ding, Junchen, et al.
Pubblicazione: (2025)
Smaller Large Language Models Can Do Moral Self-Correction
di: Liu, Guangliang, et al.
Pubblicazione: (2024)
di: Liu, Guangliang, et al.
Pubblicazione: (2024)
VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images
di: Zhou, Guanyu, et al.
Pubblicazione: (2026)
di: Zhou, Guanyu, et al.
Pubblicazione: (2026)
ProgressGym: Alignment with a Millennium of Moral Progress
di: Qiu, Tianyi, et al.
Pubblicazione: (2024)
di: Qiu, Tianyi, et al.
Pubblicazione: (2024)
Tighnari: Multi-modal Plant Species Prediction Based on Hierarchical Cross-Attention Using Graph-Based and Vision Backbone-Extracted Features
di: Liu, Haixu, et al.
Pubblicazione: (2025)
di: Liu, Haixu, et al.
Pubblicazione: (2025)
Whose Emotions and Moral Sentiments Do Language Models Reflect?
di: He, Zihao, et al.
Pubblicazione: (2024)
di: He, Zihao, et al.
Pubblicazione: (2024)
MoralCLIP: Contrastive Alignment of Vision-and-Language Representations with Moral Foundations Theory
di: Condez, Ana Carolina, et al.
Pubblicazione: (2025)
di: Condez, Ana Carolina, et al.
Pubblicazione: (2025)
Knowing But Not Doing: Convergent Morality and Divergent Action in LLMs
di: Huang, Jen-tse, et al.
Pubblicazione: (2026)
di: Huang, Jen-tse, et al.
Pubblicazione: (2026)
Do VLMs Have Bad Eyes? Diagnosing Compositional Failures via Mechanistic Interpretability
di: Aravindan, Ashwath Vaithinathan, et al.
Pubblicazione: (2025)
di: Aravindan, Ashwath Vaithinathan, et al.
Pubblicazione: (2025)
Are Language Models Sensitive to Morally Irrelevant Distractors?
di: Shaw, Andrew, et al.
Pubblicazione: (2026)
di: Shaw, Andrew, et al.
Pubblicazione: (2026)
Morally Programmed LLMs Reshape Human Morality
di: Lyu, Pengzhao, et al.
Pubblicazione: (2026)
di: Lyu, Pengzhao, et al.
Pubblicazione: (2026)
Bias Amplification Enhances Minority Group Performance
di: Li, Gaotang, et al.
Pubblicazione: (2023)
di: Li, Gaotang, et al.
Pubblicazione: (2023)
Do Emotions Influence Moral Judgment in Large Language Models?
di: Saim, Mohammad, et al.
Pubblicazione: (2026)
di: Saim, Mohammad, et al.
Pubblicazione: (2026)
Do Large Language Models Understand Morality Across Cultures?
di: Mohammadi, Hadi, et al.
Pubblicazione: (2025)
di: Mohammadi, Hadi, et al.
Pubblicazione: (2025)
The Straight and Narrow: Do LLMs Possess an Internal Moral Path?
di: Hu, Luoming, et al.
Pubblicazione: (2026)
di: Hu, Luoming, et al.
Pubblicazione: (2026)
Do Language Models Understand Morality? Towards a Robust Detection of Moral Content
di: Bulla, Luana, et al.
Pubblicazione: (2024)
di: Bulla, Luana, et al.
Pubblicazione: (2024)
That is Unacceptable: the Moral Foundations of Canceling
di: Lo, Soda Marem, et al.
Pubblicazione: (2025)
di: Lo, Soda Marem, et al.
Pubblicazione: (2025)
MOKA: Moral Knowledge Augmentation for Moral Event Extraction
di: Zhang, Xinliang Frederick, et al.
Pubblicazione: (2023)
di: Zhang, Xinliang Frederick, et al.
Pubblicazione: (2023)
MoralBERT: A Fine-Tuned Language Model for Capturing Moral Values in Social Discussions
di: Preniqi, Vjosa, et al.
Pubblicazione: (2024)
di: Preniqi, Vjosa, et al.
Pubblicazione: (2024)
Evaluating Gender Bias of LLMs in Making Morality Judgements
di: Bajaj, Divij, et al.
Pubblicazione: (2024)
di: Bajaj, Divij, et al.
Pubblicazione: (2024)
MoralBench: Moral Evaluation of LLMs
di: Ji, Jianchao, et al.
Pubblicazione: (2024)
di: Ji, Jianchao, et al.
Pubblicazione: (2024)
Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation
di: Gao, Qiyue, et al.
Pubblicazione: (2025)
di: Gao, Qiyue, et al.
Pubblicazione: (2025)
GPT-4's One-Dimensional Mapping of Morality: How the Accuracy of Country-Estimates Depends on Moral Domain
di: Strimling, Pontus, et al.
Pubblicazione: (2024)
di: Strimling, Pontus, et al.
Pubblicazione: (2024)
Decoding Multilingual Moral Preferences: Unveiling LLM's Biases Through the Moral Machine Experiment
di: Vida, Karina, et al.
Pubblicazione: (2024)
di: Vida, Karina, et al.
Pubblicazione: (2024)
The Moral Gap of Large Language Models
di: Skorski, Maciej, et al.
Pubblicazione: (2025)
di: Skorski, Maciej, et al.
Pubblicazione: (2025)
Moral Mazes in the Era of LLMs
di: Nguyen, Dang, et al.
Pubblicazione: (2026)
di: Nguyen, Dang, et al.
Pubblicazione: (2026)
The Moral Foundations Reddit Corpus
di: Trager, Jackson, et al.
Pubblicazione: (2022)
di: Trager, Jackson, et al.
Pubblicazione: (2022)
Discourse Heuristics For Paradoxically Moral Self-Correction
di: Liu, Guangliang, et al.
Pubblicazione: (2025)
di: Liu, Guangliang, et al.
Pubblicazione: (2025)
Do Consumers Accept AIs as Moral Compliance Agents?
di: Nyilasy, Greg, et al.
Pubblicazione: (2026)
di: Nyilasy, Greg, et al.
Pubblicazione: (2026)
Documenti analoghi
-
MORALISE: A Structured Benchmark for Moral Alignment in Visual Language Models
di: Lin, Xiao, et al.
Pubblicazione: (2025) -
Taming Knowledge Conflicts in Language Models
di: Li, Gaotang, et al.
Pubblicazione: (2025) -
LegalDrill: Diagnosis-Driven Synthesis for Legal Reasoning in Small Language Models
di: Li, Tianchun, et al.
Pubblicazione: (2026) -
On the Pros and Cons of Active Learning for Moral Preference Elicitation
di: Keswani, Vijay, et al.
Pubblicazione: (2024) -
Seeing but Not Believing: Probing the Disconnect Between Visual Attention and Answer Correctness in VLMs
di: Liu, Zhining, et al.
Pubblicazione: (2025)