Evaluating Moral Beliefs across LLMs through a Pluralistic Framework
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Xuelin, Zhu, Yanfei, Zhu, Shucheng, Liu, Pengyuan, Liu, Ying, Yu, Dong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Pluralistic Off-policy Evaluation and Alignment
von: Huang, Chengkai, et al.
Veröffentlicht: (2025)
von: Huang, Chengkai, et al.
Veröffentlicht: (2025)
Morality is Non-Binary: Building a Pluralist Moral Sentence Embedding Space using Contrastive Learning
von: Park, Jeongwoo, et al.
Veröffentlicht: (2024)
von: Park, Jeongwoo, et al.
Veröffentlicht: (2024)
PICACO: Pluralistic In-Context Value Alignment of LLMs via Total Correlation Optimization
von: Jiang, Han, et al.
Veröffentlicht: (2025)
von: Jiang, Han, et al.
Veröffentlicht: (2025)
A Systematic Evaluation of Preference Aggregation in Federated RLHF for Pluralistic Alignment of LLMs
von: Srewa, Mahmoud, et al.
Veröffentlicht: (2025)
von: Srewa, Mahmoud, et al.
Veröffentlicht: (2025)
ProMoral-Bench: Evaluating Prompting Strategies for Moral Reasoning and Safety in LLMs
von: Thomas, Rohan Subramanian, et al.
Veröffentlicht: (2026)
von: Thomas, Rohan Subramanian, et al.
Veröffentlicht: (2026)
THRD: A Training-Free Multi-Turn Defense Framework for Jailbreak Attacks on Large Language Models
von: Ma, Zhiqing, et al.
Veröffentlicht: (2026)
von: Ma, Zhiqing, et al.
Veröffentlicht: (2026)
MoralBench: Moral Evaluation of LLMs
von: Ji, Jianchao, et al.
Veröffentlicht: (2024)
von: Ji, Jianchao, et al.
Veröffentlicht: (2024)
MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes
von: Chiu, Yu Ying, et al.
Veröffentlicht: (2025)
von: Chiu, Yu Ying, et al.
Veröffentlicht: (2025)
The Pluralistic Moral Gap: Understanding Judgment and Value Differences between Humans and Large Language Models
von: Russo, Giuseppe, et al.
Veröffentlicht: (2025)
von: Russo, Giuseppe, et al.
Veröffentlicht: (2025)
Discerning minds or generic tutors? Evaluating instructional guidance capabilities in Socratic LLMs
von: Liu, Ying, et al.
Veröffentlicht: (2025)
von: Liu, Ying, et al.
Veröffentlicht: (2025)
A Persona-Based Evaluation Framework for Pluralistic Alignment in Generative AI
von: Karagoz, Atahan
Veröffentlicht: (2026)
von: Karagoz, Atahan
Veröffentlicht: (2026)
Exploring the Personality Traits of LLMs through Latent Features Steering
von: Yang, Shu, et al.
Veröffentlicht: (2024)
von: Yang, Shu, et al.
Veröffentlicht: (2024)
LLMs Judge Themselves: A Game-Theoretic Framework for Human-Aligned Evaluation
von: Yang, Gao, et al.
Veröffentlicht: (2025)
von: Yang, Gao, et al.
Veröffentlicht: (2025)
Fusion-Eval: Integrating Assistant Evaluators with LLMs
von: Shu, Lei, et al.
Veröffentlicht: (2023)
von: Shu, Lei, et al.
Veröffentlicht: (2023)
Layout-Aware Parsing Meets Efficient LLMs: A Unified, Scalable Framework for Resume Information Extraction and Evaluation
von: Zhu, Fanwei, et al.
Veröffentlicht: (2025)
von: Zhu, Fanwei, et al.
Veröffentlicht: (2025)
SPAGBias: Uncovering and Tracing Structured Spatial Gender Bias in Large Language Models
von: Su, Binxian, et al.
Veröffentlicht: (2026)
von: Su, Binxian, et al.
Veröffentlicht: (2026)
IslamicLegalBench: Evaluating LLMs Knowledge and Reasoning of Islamic Law Across 1,200 Years of Islamic Pluralist Legal Traditions
von: Elmahjub, Ezieddin, et al.
Veröffentlicht: (2026)
von: Elmahjub, Ezieddin, et al.
Veröffentlicht: (2026)
Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability
von: Raimondi, Bianca, et al.
Veröffentlicht: (2025)
von: Raimondi, Bianca, et al.
Veröffentlicht: (2025)
Pluralistic Alignment for Healthcare: A Role-Driven Framework
von: Zhong, Jiayou, et al.
Veröffentlicht: (2025)
von: Zhong, Jiayou, et al.
Veröffentlicht: (2025)
Hybrid Policy Distillation for LLMs
von: Zhu, Wenhong, et al.
Veröffentlicht: (2026)
von: Zhu, Wenhong, et al.
Veröffentlicht: (2026)
Vulnerability of LLMs' Stated Beliefs? LLMs Belief Resistance Check Through Strategic Persuasive Conversation Interventions
von: Huang, Fan, et al.
Veröffentlicht: (2026)
von: Huang, Fan, et al.
Veröffentlicht: (2026)
Belief in Authority: Impact of Authority in Multi-Agent Evaluation Framework
von: Choi, Junhyuk, et al.
Veröffentlicht: (2026)
von: Choi, Junhyuk, et al.
Veröffentlicht: (2026)
IDA-Bench: Evaluating LLMs on Interactive Guided Data Analysis
von: Li, Hanyu, et al.
Veröffentlicht: (2025)
von: Li, Hanyu, et al.
Veröffentlicht: (2025)
Language Models Represent Beliefs of Self and Others
von: Zhu, Wentao, et al.
Veröffentlicht: (2024)
von: Zhu, Wentao, et al.
Veröffentlicht: (2024)
CMoralEval: A Moral Evaluation Benchmark for Chinese Large Language Models
von: Yu, Linhao, et al.
Veröffentlicht: (2024)
von: Yu, Linhao, et al.
Veröffentlicht: (2024)
A Proactive EMR Assistant for Doctor-Patient Dialogue: Streaming ASR, Belief Stabilization, and Preliminary Controlled Evaluation
von: Pan, Zhenhai, et al.
Veröffentlicht: (2026)
von: Pan, Zhenhai, et al.
Veröffentlicht: (2026)
Learning to Refuse: Towards Mitigating Privacy Risks in LLMs
von: Liu, Zhenhua, et al.
Veröffentlicht: (2024)
von: Liu, Zhenhua, et al.
Veröffentlicht: (2024)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
von: Liu, Yixin, et al.
Veröffentlicht: (2025)
von: Liu, Yixin, et al.
Veröffentlicht: (2025)
Probing the Lack of Stable Internal Beliefs in LLMs
von: Luo, Yifan, et al.
Veröffentlicht: (2026)
von: Luo, Yifan, et al.
Veröffentlicht: (2026)
Rethinking the Understanding Ability across LLMs through Mutual Information
von: Wang, Shaojie, et al.
Veröffentlicht: (2025)
von: Wang, Shaojie, et al.
Veröffentlicht: (2025)
HSKBenchmark: Modeling and Benchmarking Chinese Second Language Acquisition in Large Language Models through Curriculum Tuning
von: Yang, Qihao, et al.
Veröffentlicht: (2025)
von: Yang, Qihao, et al.
Veröffentlicht: (2025)
A Roadmap to Pluralistic Alignment
von: Sorensen, Taylor, et al.
Veröffentlicht: (2024)
von: Sorensen, Taylor, et al.
Veröffentlicht: (2024)
A Survey on Transformer Context Extension: Approaches and Evaluation
von: Liu, Yijun, et al.
Veröffentlicht: (2025)
von: Liu, Yijun, et al.
Veröffentlicht: (2025)
ToolBridge: An Open-Source Dataset to Equip LLMs with External Tool Capabilities
von: Jin, Zhenchao, et al.
Veröffentlicht: (2024)
von: Jin, Zhenchao, et al.
Veröffentlicht: (2024)
Thoughts-as-Planning: Latent World Models for Chain-of-Thoughts Optimization via Reinforcement Planning
von: Liu, Dong, et al.
Veröffentlicht: (2026)
von: Liu, Dong, et al.
Veröffentlicht: (2026)
Mind the (Belief) Gap: Group Identity in the World of LLMs
von: Borah, Angana, et al.
Veröffentlicht: (2025)
von: Borah, Angana, et al.
Veröffentlicht: (2025)
Beyond Silent Letters: Amplifying LLMs in Emotion Recognition with Vocal Nuances
von: Wu, Zehui, et al.
Veröffentlicht: (2024)
von: Wu, Zehui, et al.
Veröffentlicht: (2024)
DentalBench: Benchmarking and Advancing LLMs Capability for Bilingual Dentistry Understanding
von: Zhu, Hengchuan, et al.
Veröffentlicht: (2025)
von: Zhu, Hengchuan, et al.
Veröffentlicht: (2025)
Moral Lenses, Political Coordinates: Towards Ideological Positioning of Morally Conditioned LLMs
von: Yuan, Chenchen, et al.
Veröffentlicht: (2026)
von: Yuan, Chenchen, et al.
Veröffentlicht: (2026)
Value Kaleidoscope: Engaging AI with Pluralistic Human Values, Rights, and Duties
von: Sorensen, Taylor, et al.
Veröffentlicht: (2023)
von: Sorensen, Taylor, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Pluralistic Off-policy Evaluation and Alignment
von: Huang, Chengkai, et al.
Veröffentlicht: (2025) -
Morality is Non-Binary: Building a Pluralist Moral Sentence Embedding Space using Contrastive Learning
von: Park, Jeongwoo, et al.
Veröffentlicht: (2024) -
PICACO: Pluralistic In-Context Value Alignment of LLMs via Total Correlation Optimization
von: Jiang, Han, et al.
Veröffentlicht: (2025) -
A Systematic Evaluation of Preference Aggregation in Federated RLHF for Pluralistic Alignment of LLMs
von: Srewa, Mahmoud, et al.
Veröffentlicht: (2025) -
ProMoral-Bench: Evaluating Prompting Strategies for Moral Reasoning and Safety in LLMs
von: Thomas, Rohan Subramanian, et al.
Veröffentlicht: (2026)