Evaluating Moral Beliefs across LLMs through a Pluralistic Framework
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Xuelin, Zhu, Yanfei, Zhu, Shucheng, Liu, Pengyuan, Liu, Ying, Yu, Dong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Pluralistic Off-policy Evaluation and Alignment
di: Huang, Chengkai, et al.
Pubblicazione: (2025)
di: Huang, Chengkai, et al.
Pubblicazione: (2025)
Morality is Non-Binary: Building a Pluralist Moral Sentence Embedding Space using Contrastive Learning
di: Park, Jeongwoo, et al.
Pubblicazione: (2024)
di: Park, Jeongwoo, et al.
Pubblicazione: (2024)
PICACO: Pluralistic In-Context Value Alignment of LLMs via Total Correlation Optimization
di: Jiang, Han, et al.
Pubblicazione: (2025)
di: Jiang, Han, et al.
Pubblicazione: (2025)
A Systematic Evaluation of Preference Aggregation in Federated RLHF for Pluralistic Alignment of LLMs
di: Srewa, Mahmoud, et al.
Pubblicazione: (2025)
di: Srewa, Mahmoud, et al.
Pubblicazione: (2025)
ProMoral-Bench: Evaluating Prompting Strategies for Moral Reasoning and Safety in LLMs
di: Thomas, Rohan Subramanian, et al.
Pubblicazione: (2026)
di: Thomas, Rohan Subramanian, et al.
Pubblicazione: (2026)
THRD: A Training-Free Multi-Turn Defense Framework for Jailbreak Attacks on Large Language Models
di: Ma, Zhiqing, et al.
Pubblicazione: (2026)
di: Ma, Zhiqing, et al.
Pubblicazione: (2026)
MoralBench: Moral Evaluation of LLMs
di: Ji, Jianchao, et al.
Pubblicazione: (2024)
di: Ji, Jianchao, et al.
Pubblicazione: (2024)
MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes
di: Chiu, Yu Ying, et al.
Pubblicazione: (2025)
di: Chiu, Yu Ying, et al.
Pubblicazione: (2025)
The Pluralistic Moral Gap: Understanding Judgment and Value Differences between Humans and Large Language Models
di: Russo, Giuseppe, et al.
Pubblicazione: (2025)
di: Russo, Giuseppe, et al.
Pubblicazione: (2025)
Discerning minds or generic tutors? Evaluating instructional guidance capabilities in Socratic LLMs
di: Liu, Ying, et al.
Pubblicazione: (2025)
di: Liu, Ying, et al.
Pubblicazione: (2025)
A Persona-Based Evaluation Framework for Pluralistic Alignment in Generative AI
di: Karagoz, Atahan
Pubblicazione: (2026)
di: Karagoz, Atahan
Pubblicazione: (2026)
Exploring the Personality Traits of LLMs through Latent Features Steering
di: Yang, Shu, et al.
Pubblicazione: (2024)
di: Yang, Shu, et al.
Pubblicazione: (2024)
LLMs Judge Themselves: A Game-Theoretic Framework for Human-Aligned Evaluation
di: Yang, Gao, et al.
Pubblicazione: (2025)
di: Yang, Gao, et al.
Pubblicazione: (2025)
Fusion-Eval: Integrating Assistant Evaluators with LLMs
di: Shu, Lei, et al.
Pubblicazione: (2023)
di: Shu, Lei, et al.
Pubblicazione: (2023)
Layout-Aware Parsing Meets Efficient LLMs: A Unified, Scalable Framework for Resume Information Extraction and Evaluation
di: Zhu, Fanwei, et al.
Pubblicazione: (2025)
di: Zhu, Fanwei, et al.
Pubblicazione: (2025)
SPAGBias: Uncovering and Tracing Structured Spatial Gender Bias in Large Language Models
di: Su, Binxian, et al.
Pubblicazione: (2026)
di: Su, Binxian, et al.
Pubblicazione: (2026)
IslamicLegalBench: Evaluating LLMs Knowledge and Reasoning of Islamic Law Across 1,200 Years of Islamic Pluralist Legal Traditions
di: Elmahjub, Ezieddin, et al.
Pubblicazione: (2026)
di: Elmahjub, Ezieddin, et al.
Pubblicazione: (2026)
Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability
di: Raimondi, Bianca, et al.
Pubblicazione: (2025)
di: Raimondi, Bianca, et al.
Pubblicazione: (2025)
Pluralistic Alignment for Healthcare: A Role-Driven Framework
di: Zhong, Jiayou, et al.
Pubblicazione: (2025)
di: Zhong, Jiayou, et al.
Pubblicazione: (2025)
Hybrid Policy Distillation for LLMs
di: Zhu, Wenhong, et al.
Pubblicazione: (2026)
di: Zhu, Wenhong, et al.
Pubblicazione: (2026)
Vulnerability of LLMs' Stated Beliefs? LLMs Belief Resistance Check Through Strategic Persuasive Conversation Interventions
di: Huang, Fan, et al.
Pubblicazione: (2026)
di: Huang, Fan, et al.
Pubblicazione: (2026)
Belief in Authority: Impact of Authority in Multi-Agent Evaluation Framework
di: Choi, Junhyuk, et al.
Pubblicazione: (2026)
di: Choi, Junhyuk, et al.
Pubblicazione: (2026)
IDA-Bench: Evaluating LLMs on Interactive Guided Data Analysis
di: Li, Hanyu, et al.
Pubblicazione: (2025)
di: Li, Hanyu, et al.
Pubblicazione: (2025)
Language Models Represent Beliefs of Self and Others
di: Zhu, Wentao, et al.
Pubblicazione: (2024)
di: Zhu, Wentao, et al.
Pubblicazione: (2024)
CMoralEval: A Moral Evaluation Benchmark for Chinese Large Language Models
di: Yu, Linhao, et al.
Pubblicazione: (2024)
di: Yu, Linhao, et al.
Pubblicazione: (2024)
A Proactive EMR Assistant for Doctor-Patient Dialogue: Streaming ASR, Belief Stabilization, and Preliminary Controlled Evaluation
di: Pan, Zhenhai, et al.
Pubblicazione: (2026)
di: Pan, Zhenhai, et al.
Pubblicazione: (2026)
Learning to Refuse: Towards Mitigating Privacy Risks in LLMs
di: Liu, Zhenhua, et al.
Pubblicazione: (2024)
di: Liu, Zhenhua, et al.
Pubblicazione: (2024)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
di: Liu, Yixin, et al.
Pubblicazione: (2025)
di: Liu, Yixin, et al.
Pubblicazione: (2025)
Probing the Lack of Stable Internal Beliefs in LLMs
di: Luo, Yifan, et al.
Pubblicazione: (2026)
di: Luo, Yifan, et al.
Pubblicazione: (2026)
Rethinking the Understanding Ability across LLMs through Mutual Information
di: Wang, Shaojie, et al.
Pubblicazione: (2025)
di: Wang, Shaojie, et al.
Pubblicazione: (2025)
HSKBenchmark: Modeling and Benchmarking Chinese Second Language Acquisition in Large Language Models through Curriculum Tuning
di: Yang, Qihao, et al.
Pubblicazione: (2025)
di: Yang, Qihao, et al.
Pubblicazione: (2025)
A Roadmap to Pluralistic Alignment
di: Sorensen, Taylor, et al.
Pubblicazione: (2024)
di: Sorensen, Taylor, et al.
Pubblicazione: (2024)
A Survey on Transformer Context Extension: Approaches and Evaluation
di: Liu, Yijun, et al.
Pubblicazione: (2025)
di: Liu, Yijun, et al.
Pubblicazione: (2025)
ToolBridge: An Open-Source Dataset to Equip LLMs with External Tool Capabilities
di: Jin, Zhenchao, et al.
Pubblicazione: (2024)
di: Jin, Zhenchao, et al.
Pubblicazione: (2024)
Thoughts-as-Planning: Latent World Models for Chain-of-Thoughts Optimization via Reinforcement Planning
di: Liu, Dong, et al.
Pubblicazione: (2026)
di: Liu, Dong, et al.
Pubblicazione: (2026)
Mind the (Belief) Gap: Group Identity in the World of LLMs
di: Borah, Angana, et al.
Pubblicazione: (2025)
di: Borah, Angana, et al.
Pubblicazione: (2025)
Beyond Silent Letters: Amplifying LLMs in Emotion Recognition with Vocal Nuances
di: Wu, Zehui, et al.
Pubblicazione: (2024)
di: Wu, Zehui, et al.
Pubblicazione: (2024)
DentalBench: Benchmarking and Advancing LLMs Capability for Bilingual Dentistry Understanding
di: Zhu, Hengchuan, et al.
Pubblicazione: (2025)
di: Zhu, Hengchuan, et al.
Pubblicazione: (2025)
Moral Lenses, Political Coordinates: Towards Ideological Positioning of Morally Conditioned LLMs
di: Yuan, Chenchen, et al.
Pubblicazione: (2026)
di: Yuan, Chenchen, et al.
Pubblicazione: (2026)
Value Kaleidoscope: Engaging AI with Pluralistic Human Values, Rights, and Duties
di: Sorensen, Taylor, et al.
Pubblicazione: (2023)
di: Sorensen, Taylor, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Pluralistic Off-policy Evaluation and Alignment
di: Huang, Chengkai, et al.
Pubblicazione: (2025) -
Morality is Non-Binary: Building a Pluralist Moral Sentence Embedding Space using Contrastive Learning
di: Park, Jeongwoo, et al.
Pubblicazione: (2024) -
PICACO: Pluralistic In-Context Value Alignment of LLMs via Total Correlation Optimization
di: Jiang, Han, et al.
Pubblicazione: (2025) -
A Systematic Evaluation of Preference Aggregation in Federated RLHF for Pluralistic Alignment of LLMs
di: Srewa, Mahmoud, et al.
Pubblicazione: (2025) -
ProMoral-Bench: Evaluating Prompting Strategies for Moral Reasoning and Safety in LLMs
di: Thomas, Rohan Subramanian, et al.
Pubblicazione: (2026)