The Moral Consistency Pipeline: Continuous Ethical Evaluation for Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jamshidi, Saeid, Nafi, Kawser Wazed, Dakhel, Arghavan Moradi, Shahabi, Negar, Khomh, Foutse |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Semantic Attacks on Tool-Augmented LLMs: Securing the Model Context Protocol Against Descriptor-Level Manipulation
von: Jamshidi, Saeid, et al.
Veröffentlicht: (2025)
von: Jamshidi, Saeid, et al.
Veröffentlicht: (2025)
Adversarial Moral Stress Testing of Large Language Models
von: Jamshidi, Saeid, et al.
Veröffentlicht: (2026)
von: Jamshidi, Saeid, et al.
Veröffentlicht: (2026)
Secure Tool Manifest and Digital Signing Solution for Verifiable MCP and LLM Pipelines
von: Jamshidi, Saeid, et al.
Veröffentlicht: (2026)
von: Jamshidi, Saeid, et al.
Veröffentlicht: (2026)
Tri-LLM Cooperative Federated Zero-Shot Intrusion Detection with Semantic Disagreement and Trust-Aware Aggregation
von: Jamshidi, Saeid, et al.
Veröffentlicht: (2026)
von: Jamshidi, Saeid, et al.
Veröffentlicht: (2026)
Evaluating Machine Learning-Driven Intrusion Detection Systems in IoT: Performance and Energy Consumption
von: Jamshidi, Saeid, et al.
Veröffentlicht: (2025)
von: Jamshidi, Saeid, et al.
Veröffentlicht: (2025)
Self-Adaptive Multi-Agent LLM-Based Security Pattern Selection for IoT Systems
von: Jamshidi, Saeid, et al.
Veröffentlicht: (2026)
von: Jamshidi, Saeid, et al.
Veröffentlicht: (2026)
Leveraging Machine Learning Techniques in Intrusion Detection Systems for Internet of Things
von: Jamshidi, Saeid, et al.
Veröffentlicht: (2025)
von: Jamshidi, Saeid, et al.
Veröffentlicht: (2025)
Think Fast: Real-Time IoT Intrusion Reasoning Using IDS and LLMs at the Edge Gateway
von: Jamshidi, Saeid, et al.
Veröffentlicht: (2025)
von: Jamshidi, Saeid, et al.
Veröffentlicht: (2025)
Application of Deep Reinforcement Learning for Intrusion Detection in Internet of Things: A Systematic Review
von: Jamshidi, Saeid, et al.
Veröffentlicht: (2025)
von: Jamshidi, Saeid, et al.
Veröffentlicht: (2025)
Chain of Targeted Verification Questions to Improve the Reliability of Code Generated by LLMs
von: Ngassom, Sylvain Kouemo, et al.
Veröffentlicht: (2024)
von: Ngassom, Sylvain Kouemo, et al.
Veröffentlicht: (2024)
Bugs in Large Language Models Generated Code: An Empirical Study
von: Tambon, Florian, et al.
Veröffentlicht: (2024)
von: Tambon, Florian, et al.
Veröffentlicht: (2024)
Moral Persuasion in Large Language Models: Evaluating Susceptibility and Ethical Alignment
von: Huang, Allison, et al.
Veröffentlicht: (2024)
von: Huang, Allison, et al.
Veröffentlicht: (2024)
Carbon-Aware Intrusion Detection: A Comparative Study of Supervised and Unsupervised DRL for Sustainable IoT Edge Gateways
von: Jamshidi, Saeid, et al.
Veröffentlicht: (2025)
von: Jamshidi, Saeid, et al.
Veröffentlicht: (2025)
SaGE: Evaluating Moral Consistency in Large Language Models
von: Bonagiri, Vamshi Krishna, et al.
Veröffentlicht: (2024)
von: Bonagiri, Vamshi Krishna, et al.
Veröffentlicht: (2024)
Refactoring with LLMs: Bridging Human Expertise and Machine Understanding
von: Piao, Yonnel Chen Kuang, et al.
Veröffentlicht: (2025)
von: Piao, Yonnel Chen Kuang, et al.
Veröffentlicht: (2025)
Securing Time in Energy IoT: A Clock-Dynamics-Aware Spatio-Temporal Graph Attention Network for Clock Drift Attacks and Y2K38 Failures
von: Jamshidi, Saeid, et al.
Veröffentlicht: (2026)
von: Jamshidi, Saeid, et al.
Veröffentlicht: (2026)
Multi-Agent LLM Governance for Safe Two-Timescale Reinforcement Learning in SDN-IoT Defense
von: Jamshidi, Saeid, et al.
Veröffentlicht: (2026)
von: Jamshidi, Saeid, et al.
Veröffentlicht: (2026)
"Pull or Not to Pull?'': Investigating Moral Biases in Leading Large Language Models Across Ethical Dilemmas
von: Ding, Junchen, et al.
Veröffentlicht: (2025)
von: Ding, Junchen, et al.
Veröffentlicht: (2025)
Evaluating Consistency and Reasoning Capabilities of Large Language Models
von: Saxena, Yash, et al.
Veröffentlicht: (2024)
von: Saxena, Yash, et al.
Veröffentlicht: (2024)
DCR-Consistency: Divide-Conquer-Reasoning for Consistency Evaluation and Improvement of Large Language Models
von: Cui, Wendi, et al.
Veröffentlicht: (2024)
von: Cui, Wendi, et al.
Veröffentlicht: (2024)
Chain of Simulation: A Dual-Mode Reasoning Framework for Large Language Models with Dynamic Problem Routing
von: Sheikhi, Saeid
Veröffentlicht: (2026)
von: Sheikhi, Saeid
Veröffentlicht: (2026)
CMoralEval: A Moral Evaluation Benchmark for Chinese Large Language Models
von: Yu, Linhao, et al.
Veröffentlicht: (2024)
von: Yu, Linhao, et al.
Veröffentlicht: (2024)
Cross-Examiner: Evaluating Consistency of Large Language Model-Generated Explanations
von: Villa, Danielle, et al.
Veröffentlicht: (2025)
von: Villa, Danielle, et al.
Veröffentlicht: (2025)
Firm or Fickle? Evaluating Large Language Models Consistency in Sequential Interactions
von: Li, Yubo, et al.
Veröffentlicht: (2025)
von: Li, Yubo, et al.
Veröffentlicht: (2025)
Revealing the Intrinsic Ethical Vulnerability of Aligned Large Language Models
von: Lian, Jiawei, et al.
Veröffentlicht: (2025)
von: Lian, Jiawei, et al.
Veröffentlicht: (2025)
Tracing Moral Foundations in Large Language Models
von: Yu, Chenxiao, et al.
Veröffentlicht: (2026)
von: Yu, Chenxiao, et al.
Veröffentlicht: (2026)
Cultural Bias in Large Language Models: Evaluating AI Agents through Moral Questionnaires
von: Münker, Simon
Veröffentlicht: (2025)
von: Münker, Simon
Veröffentlicht: (2025)
LP Data Pipeline: Lightweight, Purpose-driven Data Pipeline for Large Language Models
von: Kim, Yungi, et al.
Veröffentlicht: (2024)
von: Kim, Yungi, et al.
Veröffentlicht: (2024)
Ethical Reasoning and Moral Value Alignment of LLMs Depend on the Language we Prompt them in
von: Agarwal, Utkarsh, et al.
Veröffentlicht: (2024)
von: Agarwal, Utkarsh, et al.
Veröffentlicht: (2024)
Exploring and Evaluating Multimodal Knowledge Reasoning Consistency of Multimodal Large Language Models
von: Jia, Boyu, et al.
Veröffentlicht: (2025)
von: Jia, Boyu, et al.
Veröffentlicht: (2025)
Are Large Language Models Moral Hypocrites? A Study Based on Moral Foundations
von: Nunes, José Luiz, et al.
Veröffentlicht: (2024)
von: Nunes, José Luiz, et al.
Veröffentlicht: (2024)
CLLMs: Consistency Large Language Models
von: Kou, Siqi, et al.
Veröffentlicht: (2024)
von: Kou, Siqi, et al.
Veröffentlicht: (2024)
Consistency of Responses and Continuations Generated by Large Language Models on Social Media
von: Xu, Wentao, et al.
Veröffentlicht: (2025)
von: Xu, Wentao, et al.
Veröffentlicht: (2025)
A Moral Imperative: The Need for Continual Superalignment of Large Language Models
von: Puthumanaillam, Gokul, et al.
Veröffentlicht: (2024)
von: Puthumanaillam, Gokul, et al.
Veröffentlicht: (2024)
A Large Language Model Pipeline for Breast Cancer Oncology
von: Pool, Tristen, et al.
Veröffentlicht: (2024)
von: Pool, Tristen, et al.
Veröffentlicht: (2024)
EvalMORAAL: Interpretable Chain-of-Thought and LLM-as-Judge Evaluation for Moral Alignment in Large Language Models
von: Mohammadi, Hadi, et al.
Veröffentlicht: (2025)
von: Mohammadi, Hadi, et al.
Veröffentlicht: (2025)
SproutBench: A Benchmark for Safe and Ethical Large Language Models for Youth
von: Xing, Wenpeng, et al.
Veröffentlicht: (2025)
von: Xing, Wenpeng, et al.
Veröffentlicht: (2025)
Diverse Human Value Alignment for Large Language Models via Ethical Reasoning
von: Wang, Jiahao, et al.
Veröffentlicht: (2025)
von: Wang, Jiahao, et al.
Veröffentlicht: (2025)
The Only Way is Ethics: A Guide to Ethical Research with Large Language Models
von: Ungless, Eddie L., et al.
Veröffentlicht: (2024)
von: Ungless, Eddie L., et al.
Veröffentlicht: (2024)
How Do Language Models Process Ethical Instructions? Deliberation, Consistency, and Other-Recognition Across Four Models
von: Fukui, Hiroki
Veröffentlicht: (2026)
von: Fukui, Hiroki
Veröffentlicht: (2026)
Ähnliche Einträge
-
Semantic Attacks on Tool-Augmented LLMs: Securing the Model Context Protocol Against Descriptor-Level Manipulation
von: Jamshidi, Saeid, et al.
Veröffentlicht: (2025) -
Adversarial Moral Stress Testing of Large Language Models
von: Jamshidi, Saeid, et al.
Veröffentlicht: (2026) -
Secure Tool Manifest and Digital Signing Solution for Verifiable MCP and LLM Pipelines
von: Jamshidi, Saeid, et al.
Veröffentlicht: (2026) -
Tri-LLM Cooperative Federated Zero-Shot Intrusion Detection with Semantic Disagreement and Trust-Aware Aggregation
von: Jamshidi, Saeid, et al.
Veröffentlicht: (2026) -
Evaluating Machine Learning-Driven Intrusion Detection Systems in IoT: Performance and Energy Consumption
von: Jamshidi, Saeid, et al.
Veröffentlicht: (2025)