EvalMORAAL: Interpretable Chain-of-Thought and LLM-as-Judge Evaluation for Moral Alignment in Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mohammadi, Hadi, Giachanou, Anastasia, Bagheri, Robert A. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Explainability-Based Token Replacement on LLM-Generated Text
von: Mohammadi, Hadi, et al.
Veröffentlicht: (2025)
von: Mohammadi, Hadi, et al.
Veröffentlicht: (2025)
Explainability in Practice: A Survey of Explainable NLP Across Various Domains
von: Mohammadi, Hadi, et al.
Veröffentlicht: (2025)
von: Mohammadi, Hadi, et al.
Veröffentlicht: (2025)
Exploring Cultural Variations in Moral Judgments with Large Language Models
von: Mohammadi, Hadi, et al.
Veröffentlicht: (2025)
von: Mohammadi, Hadi, et al.
Veröffentlicht: (2025)
Evaluating GRPO and DPO for Faithful Chain-of-Thought Reasoning in LLMs
von: Mohammadi, Hadi, et al.
Veröffentlicht: (2025)
von: Mohammadi, Hadi, et al.
Veröffentlicht: (2025)
Large Language Models as Mirrors of Societal Moral Standards
von: Papadopoulou, Evi, et al.
Veröffentlicht: (2024)
von: Papadopoulou, Evi, et al.
Veröffentlicht: (2024)
CMoralEval: A Moral Evaluation Benchmark for Chinese Large Language Models
von: Yu, Linhao, et al.
Veröffentlicht: (2024)
von: Yu, Linhao, et al.
Veröffentlicht: (2024)
LLMs as mirrors of societal moral standards: reflection of cultural divergence and agreement across ethical topics
von: Meijer, Mijntje, et al.
Veröffentlicht: (2024)
von: Meijer, Mijntje, et al.
Veröffentlicht: (2024)
Assessing the Reliability of LLMs Annotations in the Context of Demographic Bias and Model Explanation
von: Mohammadi, Hadi, et al.
Veröffentlicht: (2025)
von: Mohammadi, Hadi, et al.
Veröffentlicht: (2025)
Gaming the Judge: Unfaithful Chain-of-Thought Can Undermine Agent Evaluation
von: Khalifa, Muhammad, et al.
Veröffentlicht: (2026)
von: Khalifa, Muhammad, et al.
Veröffentlicht: (2026)
Moral Persuasion in Large Language Models: Evaluating Susceptibility and Ethical Alignment
von: Huang, Allison, et al.
Veröffentlicht: (2024)
von: Huang, Allison, et al.
Veröffentlicht: (2024)
AdEval: Alignment-based Dynamic Evaluation to Mitigate Data Contamination in Large Language Models
von: Fan, Yang
Veröffentlicht: (2025)
von: Fan, Yang
Veröffentlicht: (2025)
AdvChain: Adversarial Chain-of-Thought Tuning for Robust Safety Alignment of Large Reasoning Models
von: Zhu, Zihao, et al.
Veröffentlicht: (2025)
von: Zhu, Zihao, et al.
Veröffentlicht: (2025)
ViLLM-Eval: A Comprehensive Evaluation Suite for Vietnamese Large Language Models
von: Nguyen, Trong-Hieu, et al.
Veröffentlicht: (2024)
von: Nguyen, Trong-Hieu, et al.
Veröffentlicht: (2024)
C2-Faith: Benchmarking LLM Judges for Causal and Coverage Faithfulness in Chain-of-Thought Reasoning
von: Mittal, Avni, et al.
Veröffentlicht: (2026)
von: Mittal, Avni, et al.
Veröffentlicht: (2026)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
von: Liu, Yixin, et al.
Veröffentlicht: (2025)
von: Liu, Yixin, et al.
Veröffentlicht: (2025)
CriticEval: Evaluating Large Language Model as Critic
von: Lan, Tian, et al.
Veröffentlicht: (2024)
von: Lan, Tian, et al.
Veröffentlicht: (2024)
Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
von: Thakur, Aman Singh, et al.
Veröffentlicht: (2024)
von: Thakur, Aman Singh, et al.
Veröffentlicht: (2024)
Applying Large Language Models and Chain-of-Thought for Automatic Scoring
von: Lee, Gyeong-Geon, et al.
Veröffentlicht: (2023)
von: Lee, Gyeong-Geon, et al.
Veröffentlicht: (2023)
YESciEval: Robust LLM-as-a-Judge for Scientific Question Answering
von: D'Souza, Jennifer, et al.
Veröffentlicht: (2025)
von: D'Souza, Jennifer, et al.
Veröffentlicht: (2025)
HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
von: Yang, Langqi, et al.
Veröffentlicht: (2025)
von: Yang, Langqi, et al.
Veröffentlicht: (2025)
How does Chain of Thought Think? Mechanistic Interpretability of Chain-of-Thought Reasoning with Sparse Autoencoding
von: Chen, Xi, et al.
Veröffentlicht: (2025)
von: Chen, Xi, et al.
Veröffentlicht: (2025)
FedCoT: Federated Chain-of-Thought Distillation for Large Language Models
von: Fan, Tao, et al.
Veröffentlicht: (2024)
von: Fan, Tao, et al.
Veröffentlicht: (2024)
Enhancing Chain-of-Thoughts Prompting with Iterative Bootstrapping in Large Language Models
von: Sun, Jiashuo, et al.
Veröffentlicht: (2023)
von: Sun, Jiashuo, et al.
Veröffentlicht: (2023)
Distributive Fairness in Large Language Models: Evaluating Alignment with Human Values
von: Hosseini, Hadi, et al.
Veröffentlicht: (2025)
von: Hosseini, Hadi, et al.
Veröffentlicht: (2025)
Do Large Language Models Understand Morality Across Cultures?
von: Mohammadi, Hadi, et al.
Veröffentlicht: (2025)
von: Mohammadi, Hadi, et al.
Veröffentlicht: (2025)
Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
von: Wu, Tianhao, et al.
Veröffentlicht: (2024)
von: Wu, Tianhao, et al.
Veröffentlicht: (2024)
When Chain-of-Thought Backfires: Evaluating Prompt Sensitivity in Medical Language Models
von: Sadanandan, Binesh, et al.
Veröffentlicht: (2026)
von: Sadanandan, Binesh, et al.
Veröffentlicht: (2026)
Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations
von: Aravindan, Ashwath Vaithinathan, et al.
Veröffentlicht: (2026)
von: Aravindan, Ashwath Vaithinathan, et al.
Veröffentlicht: (2026)
CoAT: Chain-of-Associated-Thoughts Framework for Enhancing Large Language Models Reasoning
von: Pan, Jianfeng, et al.
Veröffentlicht: (2025)
von: Pan, Jianfeng, et al.
Veröffentlicht: (2025)
Injecting Salesperson's Dialogue Strategies in Large Language Models with Chain-of-Thought Reasoning
von: Chang, Wen-Yu, et al.
Veröffentlicht: (2024)
von: Chang, Wen-Yu, et al.
Veröffentlicht: (2024)
Generalizable Chain-of-Thought Prompting in Mixed-task Scenarios with Large Language Models
von: Zou, Anni, et al.
Veröffentlicht: (2023)
von: Zou, Anni, et al.
Veröffentlicht: (2023)
The Benefits of a Concise Chain of Thought on Problem-Solving in Large Language Models
von: Renze, Matthew, et al.
Veröffentlicht: (2024)
von: Renze, Matthew, et al.
Veröffentlicht: (2024)
Improving Stance Detection by Leveraging Measurement Knowledge from Social Sciences: A Case Study of Dutch Political Tweets and Traditional Gender Role Division
von: Fang, Qixiang, et al.
Veröffentlicht: (2022)
von: Fang, Qixiang, et al.
Veröffentlicht: (2022)
Policy Frameworks for Transparent Chain-of-Thought Reasoning in Large Language Models
von: Chen, Yihang, et al.
Veröffentlicht: (2025)
von: Chen, Yihang, et al.
Veröffentlicht: (2025)
WalledEval: A Comprehensive Safety Evaluation Toolkit for Large Language Models
von: Gupta, Prannaya, et al.
Veröffentlicht: (2024)
von: Gupta, Prannaya, et al.
Veröffentlicht: (2024)
LexInstructEval: Lexical Instruction Following Evaluation for Large Language Models
von: Ren, Huimin, et al.
Veröffentlicht: (2025)
von: Ren, Huimin, et al.
Veröffentlicht: (2025)
Reasoning Beyond Chain-of-Thought: A Latent Computational Mode in Large Language Models
von: He, Zhenghao, et al.
Veröffentlicht: (2026)
von: He, Zhenghao, et al.
Veröffentlicht: (2026)
NUMCoT: Numerals and Units of Measurement in Chain-of-Thought Reasoning using Large Language Models
von: Xu, Ancheng, et al.
Veröffentlicht: (2024)
von: Xu, Ancheng, et al.
Veröffentlicht: (2024)
COFT: Counterfactual-Conformal Decoding for Fair Chain-of-Thought Reasoning in Large Language Models
von: Fayyazi, Arya, et al.
Veröffentlicht: (2026)
von: Fayyazi, Arya, et al.
Veröffentlicht: (2026)
Expediting and Elevating Large Language Model Reasoning via Hidden Chain-of-Thought Decoding
von: Liu, Tianqiao, et al.
Veröffentlicht: (2024)
von: Liu, Tianqiao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Explainability-Based Token Replacement on LLM-Generated Text
von: Mohammadi, Hadi, et al.
Veröffentlicht: (2025) -
Explainability in Practice: A Survey of Explainable NLP Across Various Domains
von: Mohammadi, Hadi, et al.
Veröffentlicht: (2025) -
Exploring Cultural Variations in Moral Judgments with Large Language Models
von: Mohammadi, Hadi, et al.
Veröffentlicht: (2025) -
Evaluating GRPO and DPO for Faithful Chain-of-Thought Reasoning in LLMs
von: Mohammadi, Hadi, et al.
Veröffentlicht: (2025) -
Large Language Models as Mirrors of Societal Moral Standards
von: Papadopoulou, Evi, et al.
Veröffentlicht: (2024)