Do VLMs Have a Moral Backbone? A Study on the Fragile Morality of Vision-Language Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Liu, Zhining, Wang, Tianyi, Lin, Xiao, Ouyang, Penghao, Li, Gaotang, Yang, Ze, Liu, Hui, Keswani, Sumit, Pardeshi, Vishwa, Zhao, Huijun, Fan, Wei, Tong, Hanghang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MORALISE: A Structured Benchmark for Moral Alignment in Visual Language Models
por: Lin, Xiao, et al.
Publicado: (2025)
por: Lin, Xiao, et al.
Publicado: (2025)
Taming Knowledge Conflicts in Language Models
por: Li, Gaotang, et al.
Publicado: (2025)
por: Li, Gaotang, et al.
Publicado: (2025)
LegalDrill: Diagnosis-Driven Synthesis for Legal Reasoning in Small Language Models
por: Li, Tianchun, et al.
Publicado: (2026)
por: Li, Tianchun, et al.
Publicado: (2026)
On the Pros and Cons of Active Learning for Moral Preference Elicitation
por: Keswani, Vijay, et al.
Publicado: (2024)
por: Keswani, Vijay, et al.
Publicado: (2024)
Seeing but Not Believing: Probing the Disconnect Between Visual Attention and Answer Correctness in VLMs
por: Liu, Zhining, et al.
Publicado: (2025)
por: Liu, Zhining, et al.
Publicado: (2025)
On The Stability of Moral Preferences: A Problem with Computational Elicitation Methods
por: Boerstler, Kyle, et al.
Publicado: (2024)
por: Boerstler, Kyle, et al.
Publicado: (2024)
The Fragility Of Moral Judgment In Large Language Models
por: van Nuenen, Tom, et al.
Publicado: (2026)
por: van Nuenen, Tom, et al.
Publicado: (2026)
SelfElicit: Your Language Model Secretly Knows Where is the Relevant Evidence
por: Liu, Zhining, et al.
Publicado: (2025)
por: Liu, Zhining, et al.
Publicado: (2025)
Moral Change or Noise? On Problems of Aligning AI With Temporally Unstable Human Feedback
por: Keswani, Vijay, et al.
Publicado: (2025)
por: Keswani, Vijay, et al.
Publicado: (2025)
Have the VLMs Lost Confidence? A Study of Sycophancy in VLMs
por: Li, Shuo, et al.
Publicado: (2024)
por: Li, Shuo, et al.
Publicado: (2024)
Can AI Model the Complexities of Human Moral Decision-Making? A Qualitative Study of Kidney Allocation Decisions
por: Keswani, Vijay, et al.
Publicado: (2025)
por: Keswani, Vijay, et al.
Publicado: (2025)
"Pull or Not to Pull?'': Investigating Moral Biases in Leading Large Language Models Across Ethical Dilemmas
por: Ding, Junchen, et al.
Publicado: (2025)
por: Ding, Junchen, et al.
Publicado: (2025)
Smaller Large Language Models Can Do Moral Self-Correction
por: Liu, Guangliang, et al.
Publicado: (2024)
por: Liu, Guangliang, et al.
Publicado: (2024)
VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images
por: Zhou, Guanyu, et al.
Publicado: (2026)
por: Zhou, Guanyu, et al.
Publicado: (2026)
ProgressGym: Alignment with a Millennium of Moral Progress
por: Qiu, Tianyi, et al.
Publicado: (2024)
por: Qiu, Tianyi, et al.
Publicado: (2024)
Tighnari: Multi-modal Plant Species Prediction Based on Hierarchical Cross-Attention Using Graph-Based and Vision Backbone-Extracted Features
por: Liu, Haixu, et al.
Publicado: (2025)
por: Liu, Haixu, et al.
Publicado: (2025)
Whose Emotions and Moral Sentiments Do Language Models Reflect?
por: He, Zihao, et al.
Publicado: (2024)
por: He, Zihao, et al.
Publicado: (2024)
MoralCLIP: Contrastive Alignment of Vision-and-Language Representations with Moral Foundations Theory
por: Condez, Ana Carolina, et al.
Publicado: (2025)
por: Condez, Ana Carolina, et al.
Publicado: (2025)
Knowing But Not Doing: Convergent Morality and Divergent Action in LLMs
por: Huang, Jen-tse, et al.
Publicado: (2026)
por: Huang, Jen-tse, et al.
Publicado: (2026)
Do VLMs Have Bad Eyes? Diagnosing Compositional Failures via Mechanistic Interpretability
por: Aravindan, Ashwath Vaithinathan, et al.
Publicado: (2025)
por: Aravindan, Ashwath Vaithinathan, et al.
Publicado: (2025)
Are Language Models Sensitive to Morally Irrelevant Distractors?
por: Shaw, Andrew, et al.
Publicado: (2026)
por: Shaw, Andrew, et al.
Publicado: (2026)
Morally Programmed LLMs Reshape Human Morality
por: Lyu, Pengzhao, et al.
Publicado: (2026)
por: Lyu, Pengzhao, et al.
Publicado: (2026)
Bias Amplification Enhances Minority Group Performance
por: Li, Gaotang, et al.
Publicado: (2023)
por: Li, Gaotang, et al.
Publicado: (2023)
Do Emotions Influence Moral Judgment in Large Language Models?
por: Saim, Mohammad, et al.
Publicado: (2026)
por: Saim, Mohammad, et al.
Publicado: (2026)
Do Large Language Models Understand Morality Across Cultures?
por: Mohammadi, Hadi, et al.
Publicado: (2025)
por: Mohammadi, Hadi, et al.
Publicado: (2025)
The Straight and Narrow: Do LLMs Possess an Internal Moral Path?
por: Hu, Luoming, et al.
Publicado: (2026)
por: Hu, Luoming, et al.
Publicado: (2026)
Do Language Models Understand Morality? Towards a Robust Detection of Moral Content
por: Bulla, Luana, et al.
Publicado: (2024)
por: Bulla, Luana, et al.
Publicado: (2024)
That is Unacceptable: the Moral Foundations of Canceling
por: Lo, Soda Marem, et al.
Publicado: (2025)
por: Lo, Soda Marem, et al.
Publicado: (2025)
MOKA: Moral Knowledge Augmentation for Moral Event Extraction
por: Zhang, Xinliang Frederick, et al.
Publicado: (2023)
por: Zhang, Xinliang Frederick, et al.
Publicado: (2023)
MoralBERT: A Fine-Tuned Language Model for Capturing Moral Values in Social Discussions
por: Preniqi, Vjosa, et al.
Publicado: (2024)
por: Preniqi, Vjosa, et al.
Publicado: (2024)
Evaluating Gender Bias of LLMs in Making Morality Judgements
por: Bajaj, Divij, et al.
Publicado: (2024)
por: Bajaj, Divij, et al.
Publicado: (2024)
MoralBench: Moral Evaluation of LLMs
por: Ji, Jianchao, et al.
Publicado: (2024)
por: Ji, Jianchao, et al.
Publicado: (2024)
Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation
por: Gao, Qiyue, et al.
Publicado: (2025)
por: Gao, Qiyue, et al.
Publicado: (2025)
GPT-4's One-Dimensional Mapping of Morality: How the Accuracy of Country-Estimates Depends on Moral Domain
por: Strimling, Pontus, et al.
Publicado: (2024)
por: Strimling, Pontus, et al.
Publicado: (2024)
Decoding Multilingual Moral Preferences: Unveiling LLM's Biases Through the Moral Machine Experiment
por: Vida, Karina, et al.
Publicado: (2024)
por: Vida, Karina, et al.
Publicado: (2024)
The Moral Gap of Large Language Models
por: Skorski, Maciej, et al.
Publicado: (2025)
por: Skorski, Maciej, et al.
Publicado: (2025)
Moral Mazes in the Era of LLMs
por: Nguyen, Dang, et al.
Publicado: (2026)
por: Nguyen, Dang, et al.
Publicado: (2026)
The Moral Foundations Reddit Corpus
por: Trager, Jackson, et al.
Publicado: (2022)
por: Trager, Jackson, et al.
Publicado: (2022)
Discourse Heuristics For Paradoxically Moral Self-Correction
por: Liu, Guangliang, et al.
Publicado: (2025)
por: Liu, Guangliang, et al.
Publicado: (2025)
Do Consumers Accept AIs as Moral Compliance Agents?
por: Nyilasy, Greg, et al.
Publicado: (2026)
por: Nyilasy, Greg, et al.
Publicado: (2026)
Ejemplares similares
-
MORALISE: A Structured Benchmark for Moral Alignment in Visual Language Models
por: Lin, Xiao, et al.
Publicado: (2025) -
Taming Knowledge Conflicts in Language Models
por: Li, Gaotang, et al.
Publicado: (2025) -
LegalDrill: Diagnosis-Driven Synthesis for Legal Reasoning in Small Language Models
por: Li, Tianchun, et al.
Publicado: (2026) -
On the Pros and Cons of Active Learning for Moral Preference Elicitation
por: Keswani, Vijay, et al.
Publicado: (2024) -
Seeing but Not Believing: Probing the Disconnect Between Visual Attention and Answer Correctness in VLMs
por: Liu, Zhining, et al.
Publicado: (2025)