Intrinsic Guardrails: How Semantic Geometry of Personality Interacts with Emergent Misalignment in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Aneja, Krishak, Mittal, Manas, Goel, Anmol, Kumaraguru, Ponnurangam, Bonagiri, Vamshi Krishna |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SaGE: Evaluating Moral Consistency in Large Language Models
by: Bonagiri, Vamshi Krishna, et al.
Published: (2024)
by: Bonagiri, Vamshi Krishna, et al.
Published: (2024)
From Human Judgements to Predictive Models: Unravelling Acceptability in Code-Mixed Sentences
by: Kodali, Prashant, et al.
Published: (2024)
by: Kodali, Prashant, et al.
Published: (2024)
COBIAS: Assessing the Contextual Reliability of Bias Benchmarks for Language Models
by: Govil, Priyanshul, et al.
Published: (2024)
by: Govil, Priyanshul, et al.
Published: (2024)
Measuring Moral Inconsistencies in Large Language Models
by: Bonagiri, Vamshi Krishna, et al.
Published: (2024)
by: Bonagiri, Vamshi Krishna, et al.
Published: (2024)
Flying Pigs, FaR and Beyond: Evaluating LLM Reasoning in Counterfactual Worlds
by: Joishy, Anish R, et al.
Published: (2025)
by: Joishy, Anish R, et al.
Published: (2025)
InSaAF: Incorporating Safety through Accuracy and Fairness | Are LLMs ready for the Indian Legal Domain?
by: Tripathi, Yogesh, et al.
Published: (2024)
by: Tripathi, Yogesh, et al.
Published: (2024)
Structured Definitions and Segmentations for Legal Reasoning in LLMs: A Study on Indian Legal Data
by: Khatri, Mann, et al.
Published: (2025)
by: Khatri, Mann, et al.
Published: (2025)
Check Yourself Before You Wreck Yourself: Selectively Quitting Improves LLM Agent Safety
by: Bonagiri, Vamshi Krishna, et al.
Published: (2025)
by: Bonagiri, Vamshi Krishna, et al.
Published: (2025)
Shadow Unlearning: A Neuro-Semantic Approach to Fidelity-Preserving Faceless Forgetting in LLMs
by: P, Dinesh Srivasthav, et al.
Published: (2026)
by: P, Dinesh Srivasthav, et al.
Published: (2026)
Do LLMs Adhere to Label Definitions? Examining Their Receptivity to External Label Definitions
by: Mohammadi, Seyedali, et al.
Published: (2025)
by: Mohammadi, Seyedali, et al.
Published: (2025)
Wu's Method can Boost Symbolic AI to Rival Silver Medalists and AlphaGeometry to Outperform Gold Medalists at IMO Geometry
by: Sinha, Shiven, et al.
Published: (2024)
by: Sinha, Shiven, et al.
Published: (2024)
PrivacyBench: A Conversational Benchmark for Evaluating Privacy in Personalized AI
by: Mukhopadhyay, Srija, et al.
Published: (2025)
by: Mukhopadhyay, Srija, et al.
Published: (2025)
Semantic Containment as a Fundamental Property of Emergent Misalignment
by: Saxena, Rohan
Published: (2026)
by: Saxena, Rohan
Published: (2026)
SEMMA: A Semantic Aware Knowledge Graph Foundation Model
by: Arun, Arvindh, et al.
Published: (2025)
by: Arun, Arvindh, et al.
Published: (2025)
HLDC: Hindi Legal Documents Corpus
by: Kapoor, Arnav, et al.
Published: (2022)
by: Kapoor, Arnav, et al.
Published: (2022)
LLMs Deceive Unintentionally: Emergent Misalignment in Dishonesty from Misaligned Samples to Biased Human-AI Interactions
by: Hu, Xuhao, et al.
Published: (2025)
by: Hu, Xuhao, et al.
Published: (2025)
Multilingual Coreference Resolution in Low-resource South Asian Languages
by: Mishra, Ritwik, et al.
Published: (2024)
by: Mishra, Ritwik, et al.
Published: (2024)
Emergent Misalignment is Easy, Narrow Misalignment is Hard
by: Soligo, Anna, et al.
Published: (2026)
by: Soligo, Anna, et al.
Published: (2026)
Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs
by: Giordani, Jeremiah
Published: (2025)
by: Giordani, Jeremiah
Published: (2025)
ImplicitBBQ: Benchmarking Implicit Bias in Large Language Models through Characteristic Based Cues
by: Vedula, Bhaskara Hanuma, et al.
Published: (2026)
by: Vedula, Bhaskara Hanuma, et al.
Published: (2026)
The Devil in the Details: Emergent Misalignment, Format and Coherence in Open-Weights LLMs
by: Dickson, Craig
Published: (2025)
by: Dickson, Craig
Published: (2025)
Multilingual Non-Factoid Question Answering with Answer Paragraph Selection
by: Mishra, Ritwik, et al.
Published: (2024)
by: Mishra, Ritwik, et al.
Published: (2024)
Just KIDDIN: Knowledge Infusion and Distillation for Detection of INdecent Memes
by: Garg, Rahul, et al.
Published: (2024)
by: Garg, Rahul, et al.
Published: (2024)
Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
by: Betley, Jan, et al.
Published: (2025)
by: Betley, Jan, et al.
Published: (2025)
Great Models Think Alike and this Undermines AI Oversight
by: Goel, Shashwat, et al.
Published: (2025)
by: Goel, Shashwat, et al.
Published: (2025)
KnowledgePrompts: Exploring the Abilities of Large Language Models to Solve Proportional Analogies via Knowledge-Enhanced Prompting
by: Wijesiriwardene, Thilini, et al.
Published: (2024)
by: Wijesiriwardene, Thilini, et al.
Published: (2024)
Manifold-based Sampling for In-Context Hallucination Detection in Large Language Models
by: Vamshi, Bodla Krishna, et al.
Published: (2026)
by: Vamshi, Bodla Krishna, et al.
Published: (2026)
Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models
by: Chua, James, et al.
Published: (2025)
by: Chua, James, et al.
Published: (2025)
Persona-Model Collapse in Emergent Misalignment
by: Costa, Davi Bastos, et al.
Published: (2026)
by: Costa, Davi Bastos, et al.
Published: (2026)
Interpretable Question Answering with Knowledge Graphs
by: Aneja, Kartikeya, et al.
Published: (2025)
by: Aneja, Kartikeya, et al.
Published: (2025)
Are Models Trained on Indian Legal Data Fair?
by: Girhepuje, Sahil, et al.
Published: (2023)
by: Girhepuje, Sahil, et al.
Published: (2023)
Emergent and Subliminal Misalignment Through the Lens of Data-Mediated Transfer
by: Askin, Baris, et al.
Published: (2026)
by: Askin, Baris, et al.
Published: (2026)
Misaligned by Reward: Socially Undesirable Preferences in LLMs
by: Ghazaryan, Gayane, et al.
Published: (2026)
by: Ghazaryan, Gayane, et al.
Published: (2026)
Eliciting and Analyzing Emergent Misalignment in State-of-the-Art Large Language Models
by: Panpatil, Siddhant, et al.
Published: (2025)
by: Panpatil, Siddhant, et al.
Published: (2025)
Mechanistic Steering of LLMs Reveals Layer-wise Feature Vulnerabilities in Adversarial Settings
by: Das, Nilanjana, et al.
Published: (2026)
by: Das, Nilanjana, et al.
Published: (2026)
RAG Makes Guardrails Unsafe? Investigating Robustness of Guardrails under RAG-style Contexts
by: She, Yining, et al.
Published: (2025)
by: She, Yining, et al.
Published: (2025)
I Can't Believe It's Corrupt: Evaluating Corruption in Multi-Agent Governance Systems
by: P, Vedanta S, et al.
Published: (2026)
by: P, Vedanta S, et al.
Published: (2026)
Semantic Integrity Constraints: Declarative Guardrails for AI-Augmented Data Processing Systems
by: Lee, Alexander W., et al.
Published: (2025)
by: Lee, Alexander W., et al.
Published: (2025)
Corrective Machine Unlearning
by: Goel, Shashwat, et al.
Published: (2024)
by: Goel, Shashwat, et al.
Published: (2024)
Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment
by: Krishna, Kundan, et al.
Published: (2025)
by: Krishna, Kundan, et al.
Published: (2025)
Similar Items
-
SaGE: Evaluating Moral Consistency in Large Language Models
by: Bonagiri, Vamshi Krishna, et al.
Published: (2024) -
From Human Judgements to Predictive Models: Unravelling Acceptability in Code-Mixed Sentences
by: Kodali, Prashant, et al.
Published: (2024) -
COBIAS: Assessing the Contextual Reliability of Bias Benchmarks for Language Models
by: Govil, Priyanshul, et al.
Published: (2024) -
Measuring Moral Inconsistencies in Large Language Models
by: Bonagiri, Vamshi Krishna, et al.
Published: (2024) -
Flying Pigs, FaR and Beyond: Evaluating LLM Reasoning in Counterfactual Worlds
by: Joishy, Anish R, et al.
Published: (2025)