Do Language Models Understand Morality? Towards a Robust Detection of Moral Content
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bulla, Luana, Gangemi, Aldo, Mongiovì, Misael |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Leveraging Large Language Models for Accurate Sign Language Translation in Low-Resource Scenarios
von: Bulla, Luana, et al.
Veröffentlicht: (2025)
von: Bulla, Luana, et al.
Veröffentlicht: (2025)
Enhancing Retrieval-Augmented Generation with Entity Linking for Educational Platforms
von: Granata, Francesco, et al.
Veröffentlicht: (2025)
von: Granata, Francesco, et al.
Veröffentlicht: (2025)
Explainable Moral Values: a neuro-symbolic approach to value classification
von: Lazzari, Nicolas, et al.
Veröffentlicht: (2024)
von: Lazzari, Nicolas, et al.
Veröffentlicht: (2024)
Morality is Contextual: Learning Interpretable Moral Contexts from Human Data with Probabilistic Clustering and Large Language Models
von: Morlat, Geoffroy, et al.
Veröffentlicht: (2025)
von: Morlat, Geoffroy, et al.
Veröffentlicht: (2025)
Do VLMs Have a Moral Backbone? A Study on the Fragile Morality of Vision-Language Models
von: Liu, Zhining, et al.
Veröffentlicht: (2026)
von: Liu, Zhining, et al.
Veröffentlicht: (2026)
A Moral Imperative: The Need for Continual Superalignment of Large Language Models
von: Puthumanaillam, Gokul, et al.
Veröffentlicht: (2024)
von: Puthumanaillam, Gokul, et al.
Veröffentlicht: (2024)
Does Cross-Cultural Alignment Change the Commonsense Morality of Language Models?
von: Jinnai, Yuu
Veröffentlicht: (2024)
von: Jinnai, Yuu
Veröffentlicht: (2024)
Language over Content: Tracing Cultural Understanding in Multilingual Large Language Models
von: Cho, Seungho, et al.
Veröffentlicht: (2025)
von: Cho, Seungho, et al.
Veröffentlicht: (2025)
GreedLlama: Performance of Financial Value-Aligned Large Language Models in Moral Reasoning
von: Yu, Jeffy, et al.
Veröffentlicht: (2024)
von: Yu, Jeffy, et al.
Veröffentlicht: (2024)
Transformers and Slot Encoding for Sample Efficient Physical World Modelling
von: Petri, Francesco, et al.
Veröffentlicht: (2024)
von: Petri, Francesco, et al.
Veröffentlicht: (2024)
Deep Learning Detection Method for Large Language Models-Generated Scientific Content
von: Alhijawi, Bushra, et al.
Veröffentlicht: (2024)
von: Alhijawi, Bushra, et al.
Veröffentlicht: (2024)
Logic Augmented Generation
von: Gangemi, Aldo, et al.
Veröffentlicht: (2024)
von: Gangemi, Aldo, et al.
Veröffentlicht: (2024)
AutoDetect: Towards a Unified Framework for Automated Weakness Detection in Large Language Models
von: Cheng, Jiale, et al.
Veröffentlicht: (2024)
von: Cheng, Jiale, et al.
Veröffentlicht: (2024)
Towards Understanding Multi-Round Large Language Model Reasoning: Approximability, Learnability and Generalizability
von: Xu, Chenhui, et al.
Veröffentlicht: (2025)
von: Xu, Chenhui, et al.
Veröffentlicht: (2025)
Towards Robust Alignment of Language Models: Distributionally Robustifying Direct Preference Optimization
von: Wu, Junkang, et al.
Veröffentlicht: (2024)
von: Wu, Junkang, et al.
Veröffentlicht: (2024)
That's Deprecated! Understanding, Detecting, and Steering Knowledge Conflicts in Language Models for Code Generation
von: Bae, Jaesung, et al.
Veröffentlicht: (2025)
von: Bae, Jaesung, et al.
Veröffentlicht: (2025)
TF1-EN-3M: Three Million Synthetic Moral Fables for Training Small, Open Language Models
von: Nadas, Mihai, et al.
Veröffentlicht: (2025)
von: Nadas, Mihai, et al.
Veröffentlicht: (2025)
Understanding Moral Reasoning Trajectories in Large Language Models: Toward Probing-Based Explainability
von: Huang, Fan, et al.
Veröffentlicht: (2026)
von: Huang, Fan, et al.
Veröffentlicht: (2026)
Towards Understanding the Robustness of Sparse Autoencoders
von: Saiyed, Ahson, et al.
Veröffentlicht: (2026)
von: Saiyed, Ahson, et al.
Veröffentlicht: (2026)
(How) Do Language Models Track State?
von: Li, Belinda Z., et al.
Veröffentlicht: (2025)
von: Li, Belinda Z., et al.
Veröffentlicht: (2025)
Language Models Use Trigonometry to Do Addition
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2025)
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2025)
MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes
von: Chiu, Yu Ying, et al.
Veröffentlicht: (2025)
von: Chiu, Yu Ying, et al.
Veröffentlicht: (2025)
ProgressGym: Alignment with a Millennium of Moral Progress
von: Qiu, Tianyi, et al.
Veröffentlicht: (2024)
von: Qiu, Tianyi, et al.
Veröffentlicht: (2024)
Towards Understanding Sycophancy in Language Models
von: Sharma, Mrinank, et al.
Veröffentlicht: (2023)
von: Sharma, Mrinank, et al.
Veröffentlicht: (2023)
Understanding Understanding: A Pragmatic Framework Motivated by Large Language Models
von: Leyton-Brown, Kevin, et al.
Veröffentlicht: (2024)
von: Leyton-Brown, Kevin, et al.
Veröffentlicht: (2024)
Towards Interpretable Hate Speech Detection using Large Language Model-extracted Rationales
von: Nirmal, Ayushi, et al.
Veröffentlicht: (2024)
von: Nirmal, Ayushi, et al.
Veröffentlicht: (2024)
Language Models Do Hard Arithmetic Tasks Easily and Hardly Do Easy Arithmetic Tasks
von: Gambardella, Andrew, et al.
Veröffentlicht: (2024)
von: Gambardella, Andrew, et al.
Veröffentlicht: (2024)
On the Robustness of Reward Models for Language Model Alignment
von: Hong, Jiwoo, et al.
Veröffentlicht: (2025)
von: Hong, Jiwoo, et al.
Veröffentlicht: (2025)
Understanding and Mitigating Tokenization Bias in Language Models
von: Phan, Buu, et al.
Veröffentlicht: (2024)
von: Phan, Buu, et al.
Veröffentlicht: (2024)
Emergent Semantic Role Understanding in Language Models
von: Griffiths, Carla, et al.
Veröffentlicht: (2026)
von: Griffiths, Carla, et al.
Veröffentlicht: (2026)
Understanding Subword Compositionality of Large Language Models
von: Peng, Qiwei, et al.
Veröffentlicht: (2025)
von: Peng, Qiwei, et al.
Veröffentlicht: (2025)
Semi-Supervised Learning for Large Language Models Safety and Content Moderation
von: Dinuta, Eduard Stefan, et al.
Veröffentlicht: (2025)
von: Dinuta, Eduard Stefan, et al.
Veröffentlicht: (2025)
Do Large Language Models Show Biases in Causal Learning?
von: Carro, Maria Victoria, et al.
Veröffentlicht: (2024)
von: Carro, Maria Victoria, et al.
Veröffentlicht: (2024)
Why Larger Language Models Do In-context Learning Differently?
von: Shi, Zhenmei, et al.
Veröffentlicht: (2024)
von: Shi, Zhenmei, et al.
Veröffentlicht: (2024)
Do Large Language Models Know How Much They Know?
von: Prato, Gabriele, et al.
Veröffentlicht: (2025)
von: Prato, Gabriele, et al.
Veröffentlicht: (2025)
Toward Understanding BERT-Like Pre-Training for DNA Foundation Models
von: Liang, Chaoqi, et al.
Veröffentlicht: (2023)
von: Liang, Chaoqi, et al.
Veröffentlicht: (2023)
GUNDAM: Aligning Large Language Models with Graph Understanding
von: Ouyang, Sheng, et al.
Veröffentlicht: (2024)
von: Ouyang, Sheng, et al.
Veröffentlicht: (2024)
Understanding and Accelerating the Training of Masked Diffusion Language Models
von: Hong, Chunsan, et al.
Veröffentlicht: (2026)
von: Hong, Chunsan, et al.
Veröffentlicht: (2026)
Do I Know This Entity? Knowledge Awareness and Hallucinations in Language Models
von: Ferrando, Javier, et al.
Veröffentlicht: (2024)
von: Ferrando, Javier, et al.
Veröffentlicht: (2024)
Do Retrieval-Augmented Language Models Adapt to Varying User Needs?
von: Wu, Peilin, et al.
Veröffentlicht: (2025)
von: Wu, Peilin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Leveraging Large Language Models for Accurate Sign Language Translation in Low-Resource Scenarios
von: Bulla, Luana, et al.
Veröffentlicht: (2025) -
Enhancing Retrieval-Augmented Generation with Entity Linking for Educational Platforms
von: Granata, Francesco, et al.
Veröffentlicht: (2025) -
Explainable Moral Values: a neuro-symbolic approach to value classification
von: Lazzari, Nicolas, et al.
Veröffentlicht: (2024) -
Morality is Contextual: Learning Interpretable Moral Contexts from Human Data with Probabilistic Clustering and Large Language Models
von: Morlat, Geoffroy, et al.
Veröffentlicht: (2025) -
Do VLMs Have a Moral Backbone? A Study on the Fragile Morality of Vision-Language Models
von: Liu, Zhining, et al.
Veröffentlicht: (2026)