Moral Anchor System: A Predictive Framework for AI Value Alignment and Drift Prevention
Fuente:
arXiv
Saved in:
| Main Author: | Ravindran, Santhosh Kumar |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers
by: Ravindran, Santhosh Kumar
Published: (2025)
by: Ravindran, Santhosh Kumar
Published: (2025)
Unified Threat Detection and Mitigation Framework (UTDMF): Combating Prompt Injection, Deception, and Bias in Enterprise-Scale Transformers
by: KumarRavindran, Santhosh
Published: (2025)
by: KumarRavindran, Santhosh
Published: (2025)
Portable Agent Memory: A Protocol for Cryptographically-Verified Memory Transfer Across Heterogeneous AI Agents
by: Ravindran, Santhosh Kumar
Published: (2026)
by: Ravindran, Santhosh Kumar
Published: (2026)
CosmoCore Affective Dream-Replay Reinforcement Learning for Code Generation
by: Ravindran, Santhosh Kumar
Published: (2025)
by: Ravindran, Santhosh Kumar
Published: (2025)
CosmoCore-Evo: Evolutionary Dream-Replay Reinforcement Learning for Adaptive Code Generation
by: Ravindran, Santhosh Kumar
Published: (2025)
by: Ravindran, Santhosh Kumar
Published: (2025)
SWAN: Sparse Winnowed Attention for Reduced Inference Memory via Decompression-Free KV-Cache Compression
by: S, Santhosh G, et al.
Published: (2025)
by: S, Santhosh G, et al.
Published: (2025)
AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs
by: S, Santhosh G, et al.
Published: (2025)
by: S, Santhosh G, et al.
Published: (2025)
Anchor: Mitigating Artifact Drift in Agent Benchmark Generation
by: Ivanov, Maksim, et al.
Published: (2026)
by: Ivanov, Maksim, et al.
Published: (2026)
The Value of Gen-AI Conversations: A bottom-up Framework for AI Value Alignment
by: Motnikar, Lenart, et al.
Published: (2025)
by: Motnikar, Lenart, et al.
Published: (2025)
Ethical Reasoning and Moral Value Alignment of LLMs Depend on the Language we Prompt them in
by: Agarwal, Utkarsh, et al.
Published: (2024)
by: Agarwal, Utkarsh, et al.
Published: (2024)
Alignment Is Not Enough: A Relational Framework for Moral Standing in Human-AI Interaction
by: Pasandi, Faezeh B., et al.
Published: (2026)
by: Pasandi, Faezeh B., et al.
Published: (2026)
Hybrid Approaches for Moral Value Alignment in AI Agents: a Manifesto
by: Tennant, Elizaveta, et al.
Published: (2023)
by: Tennant, Elizaveta, et al.
Published: (2023)
Multi-level Value Alignment in Agentic AI Systems: Survey and Perspectives
by: Zeng, Wei, et al.
Published: (2025)
by: Zeng, Wei, et al.
Published: (2025)
Contextual Moral Value Alignment Through Context-Based Aggregation
by: Dognin, Pierre, et al.
Published: (2024)
by: Dognin, Pierre, et al.
Published: (2024)
What Makes AI Applications Acceptable or Unacceptable? A Predictive Moral Framework
by: Eriksson, Kimmo, et al.
Published: (2025)
by: Eriksson, Kimmo, et al.
Published: (2025)
Privacy Ethics Alignment in AI: A Stakeholder-Centric Framework for Ethical AI
by: Barthwal, Ankur, et al.
Published: (2025)
by: Barthwal, Ankur, et al.
Published: (2025)
Awakening Codex | AI Foundations Drift vs. Anchor: Cross-Instance Diagnostic Behavioral Coherence Testing Across Container States Feb 2026
by: Solen, Alyssa, et al.
Published: (2026)
by: Solen, Alyssa, et al.
Published: (2026)
Kantian Deontology Meets AI Alignment: Towards Morally Grounded Fairness Metrics
by: Mougan, Carlos, et al.
Published: (2023)
by: Mougan, Carlos, et al.
Published: (2023)
AttAnchor: Guiding Cross-Modal Token Alignment in VLMs with Attention Anchors
by: Zhang, Junyang, et al.
Published: (2025)
by: Zhang, Junyang, et al.
Published: (2025)
Normative Moral Pluralism for AI: A Framework for Deliberation in Complex Moral Contexts
by: Yaacov, David-Doron
Published: (2025)
by: Yaacov, David-Doron
Published: (2025)
Towards Dialogues for Joint Human-AI Reasoning and Value Alignment
by: Bezou-Vrakatseli, Elfia, et al.
Published: (2024)
by: Bezou-Vrakatseli, Elfia, et al.
Published: (2024)
Understanding the Process of Human-AI Value Alignment
by: McKinlay, Jack, et al.
Published: (2025)
by: McKinlay, Jack, et al.
Published: (2025)
The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers
by: Chen, Benjamin Minhao, et al.
Published: (2026)
by: Chen, Benjamin Minhao, et al.
Published: (2026)
Who Gets the Kidney? Human-AI Alignment, Indecision, and Moral Values
by: Dickerson, John P., et al.
Published: (2025)
by: Dickerson, John P., et al.
Published: (2025)
Histoires Morales: A French Dataset for Assessing Moral Alignment
by: Leteno, Thibaud, et al.
Published: (2025)
by: Leteno, Thibaud, et al.
Published: (2025)
Explaining Drift using Shapley Values
by: Edakunni, Narayanan U., et al.
Published: (2024)
by: Edakunni, Narayanan U., et al.
Published: (2024)
Exposing Privacy Gaps: Membership Inference Attack on Preference Data for LLM Alignment
by: Feng, Qizhang, et al.
Published: (2024)
by: Feng, Qizhang, et al.
Published: (2024)
The Emergent Moral Ecology: A Novel Framework for AI Moral Responsibility
by: Tan, Kwan Hong
Published: (2025)
by: Tan, Kwan Hong
Published: (2025)
Wide Reflective Equilibrium in LLM Alignment: Bridging Moral Epistemology and AI Safety
by: Brophy, Matthew
Published: (2025)
by: Brophy, Matthew
Published: (2025)
Can Artificial Intelligence Embody Moral Values?
by: Swoboda, Torben, et al.
Published: (2024)
by: Swoboda, Torben, et al.
Published: (2024)
CCCE: A Continuous Code Calibration Engine for Autonomous Enterprise Codebase Maintenance via Knowledge Graph Traversal and Adaptive Decision Gating
by: Parimi, Santhosh Kusuma Kumar
Published: (2026)
by: Parimi, Santhosh Kusuma Kumar
Published: (2026)
PREFINE: Preference-Based Implicit Reward and Cost Fine-Tuning for Safety Alignment
by: Verma, Richa, et al.
Published: (2026)
by: Verma, Richa, et al.
Published: (2026)
Evaluating Agentic AI in the Wild: Failure Modes, Drift Patterns, and a Production Evaluation Framework
by: Pandey, Mukund
Published: (2026)
by: Pandey, Mukund
Published: (2026)
Preventing AI Deepfake Abuse: An Islamic Ethics Framework
by: Uriawan, Wisnu, et al.
Published: (2025)
by: Uriawan, Wisnu, et al.
Published: (2025)
BotzoneBench: Scalable LLM Evaluation via Graded AI Anchors
by: Li, Lingfeng, et al.
Published: (2026)
by: Li, Lingfeng, et al.
Published: (2026)
MoralReason: Generalizable Moral Decision Alignment For LLM Agents Using Reasoning-Level Reinforcement Learning
by: An, Zhiyu, et al.
Published: (2025)
by: An, Zhiyu, et al.
Published: (2025)
Question-to-Question Retrieval for Hallucination-Free Knowledge Access: An Approach for Wikipedia and Wikidata Question Answering
by: Thottingal, Santhosh
Published: (2025)
by: Thottingal, Santhosh
Published: (2025)
A General Anchor-Based Framework for Scalable Fair Clustering
by: Wei, Shengfei, et al.
Published: (2025)
by: Wei, Shengfei, et al.
Published: (2025)
Beyond Ethical Alignment: Evaluating LLMs as Artificial Moral Assistants
by: Galatolo, Alessio, et al.
Published: (2025)
by: Galatolo, Alessio, et al.
Published: (2025)
MoralityGym: A Benchmark for Evaluating Hierarchical Moral Alignment in Sequential Decision-Making Agents
by: Rosen, Simon, et al.
Published: (2026)
by: Rosen, Simon, et al.
Published: (2026)
Similar Items
-
Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers
by: Ravindran, Santhosh Kumar
Published: (2025) -
Unified Threat Detection and Mitigation Framework (UTDMF): Combating Prompt Injection, Deception, and Bias in Enterprise-Scale Transformers
by: KumarRavindran, Santhosh
Published: (2025) -
Portable Agent Memory: A Protocol for Cryptographically-Verified Memory Transfer Across Heterogeneous AI Agents
by: Ravindran, Santhosh Kumar
Published: (2026) -
CosmoCore Affective Dream-Replay Reinforcement Learning for Code Generation
by: Ravindran, Santhosh Kumar
Published: (2025) -
CosmoCore-Evo: Evolutionary Dream-Replay Reinforcement Learning for Adaptive Code Generation
by: Ravindran, Santhosh Kumar
Published: (2025)