Moral Sensitivity in LLMs: A Tiered Evaluation of Contextual Bias via Behavioral Profiling and Mechanistic Interpretability
Fuente:
arXiv
Guardado en:
| Autores principales: | Aggarwal, Yash, Gorti, Atmika, Jain, Vinija, Chadha, Aman, Thirunarayan, Krishnaprasad, Gaur, Manas |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Unboxing Occupational Bias: Grounded Debiasing of LLMs with U.S. Labor Data
por: Gorti, Atmika, et al.
Publicado: (2024)
por: Gorti, Atmika, et al.
Publicado: (2024)
Mental Health Equity in LLMs: Leveraging Multi-Hop Question Answering to Detect Amplified and Silenced Perspectives
por: Haider, Batool, et al.
Publicado: (2025)
por: Haider, Batool, et al.
Publicado: (2025)
The Reasoning Trap -- Logical Reasoning as a Mechanistic Pathway to Situational Awareness
por: Sahoo, Subramanyam, et al.
Publicado: (2026)
por: Sahoo, Subramanyam, et al.
Publicado: (2026)
Born With a Silver Spoon? Investigating Socioeconomic Bias in Large Language Models
por: Singh, Smriti, et al.
Publicado: (2024)
por: Singh, Smriti, et al.
Publicado: (2024)
Dial E for Ethical Enforcement: institutional VETO power as a governance primitive
por: Sahoo, Subramanyam, et al.
Publicado: (2026)
por: Sahoo, Subramanyam, et al.
Publicado: (2026)
TRACEALIGN -- Tracing the Drift: Attributing Alignment Failures to Training-Time Belief Sources in LLMs
por: Das, Amitava, et al.
Publicado: (2025)
por: Das, Amitava, et al.
Publicado: (2025)
COBIAS: Assessing the Contextual Reliability of Bias Benchmarks for Language Models
por: Govil, Priyanshul, et al.
Publicado: (2024)
por: Govil, Priyanshul, et al.
Publicado: (2024)
Flying Pigs, FaR and Beyond: Evaluating LLM Reasoning in Counterfactual Worlds
por: Joishy, Anish R, et al.
Publicado: (2025)
por: Joishy, Anish R, et al.
Publicado: (2025)
Reasoning or Rhetoric? An Empirical Analysis of Moral Reasoning Explanations in Large Language Models
por: Kasat, Aryan, et al.
Publicado: (2026)
por: Kasat, Aryan, et al.
Publicado: (2026)
From Prejudice to Parity: A New Approach to Debiasing Large Language Model Word Embeddings
por: Rakshit, Aishik, et al.
Publicado: (2024)
por: Rakshit, Aishik, et al.
Publicado: (2024)
Guiding Vision-Language Model Selection for Visual Question-Answering Across Tasks, Domains, and Knowledge Types
por: Sinha, Neelabh, et al.
Publicado: (2024)
por: Sinha, Neelabh, et al.
Publicado: (2024)
Are Small Language Models Ready to Compete with Large Language Models for Practical Applications?
por: Sinha, Neelabh, et al.
Publicado: (2024)
por: Sinha, Neelabh, et al.
Publicado: (2024)
D-STEER - Preference Alignment Techniques Learn to Behave, not to Believe -- Beneath the Surface, DPO as Steering Vector Perturbation in Activation Space
por: Raina, Samarth, et al.
Publicado: (2025)
por: Raina, Samarth, et al.
Publicado: (2025)
Personality Shapes Gender Bias in Persona-Conditioned LLM Narratives Across English and Hindi: An Empirical Investigation
por: Kumar, Tanay, et al.
Publicado: (2026)
por: Kumar, Tanay, et al.
Publicado: (2026)
AlignGuard-LoRA: Alignment-Preserving Fine-Tuning via Fisher-Guided Decomposition and Riemannian-Geodesic Collision Regularization
por: Das, Amitava, et al.
Publicado: (2025)
por: Das, Amitava, et al.
Publicado: (2025)
MAAT: Multi-phase Adapter-Aware Targeted Unlearning
por: Yagnik, Suryash, et al.
Publicado: (2026)
por: Yagnik, Suryash, et al.
Publicado: (2026)
Neural FOXP2 -- Language Specific Neuron Steering for Targeted Language Improvement in LLMs
por: Saha, Anusa, et al.
Publicado: (2026)
por: Saha, Anusa, et al.
Publicado: (2026)
Towards Robust Evaluation of Unlearning in LLMs via Data Transformations
por: Joshi, Abhinav, et al.
Publicado: (2024)
por: Joshi, Abhinav, et al.
Publicado: (2024)
LLMsAgainstHate @ NLU of Devanagari Script Languages 2025: Hate Speech Detection and Target Identification in Devanagari Languages via Parameter Efficient Fine-Tuning of LLMs
por: Sidibomma, Rushendra, et al.
Publicado: (2024)
por: Sidibomma, Rushendra, et al.
Publicado: (2024)
When Shallow Wins: Silent Failures and the Depth-Accuracy Paradox in Latent Reasoning
por: Sahoo, Subramanyam, et al.
Publicado: (2026)
por: Sahoo, Subramanyam, et al.
Publicado: (2026)
SAHOO: Safeguarded Alignment for High-Order Optimization Objectives in Recursive Self-Improvement
por: Sahoo, Subramanyam, et al.
Publicado: (2026)
por: Sahoo, Subramanyam, et al.
Publicado: (2026)
I Can't Believe It's Not Robust: Catastrophic Collapse of Safety Classifiers under Embedding Drift
por: Sahoo, Subramanyam, et al.
Publicado: (2026)
por: Sahoo, Subramanyam, et al.
Publicado: (2026)
Position: The Complexity of Perfect AI Alignment -- Formalizing the RLHF Trilemma
por: Sahoo, Subramanyam, et al.
Publicado: (2025)
por: Sahoo, Subramanyam, et al.
Publicado: (2025)
Exploring the Impact of Large Language Models on Recommender Systems: An Extensive Review
por: Vats, Arpita, et al.
Publicado: (2024)
por: Vats, Arpita, et al.
Publicado: (2024)
SleepWalk: A Three-Tier Benchmark for Stress-Testing Instruction-Guided Vision-Language Navigation
por: Rawal, Niyati, et al.
Publicado: (2026)
por: Rawal, Niyati, et al.
Publicado: (2026)
Human-Readable Adversarial Prompts: An Investigation into LLM Vulnerabilities Using Situational Context
por: Das, Nilanjana, et al.
Publicado: (2024)
por: Das, Nilanjana, et al.
Publicado: (2024)
Mechanistic Steering of LLMs Reveals Layer-wise Feature Vulnerabilities in Adversarial Settings
por: Das, Nilanjana, et al.
Publicado: (2026)
por: Das, Nilanjana, et al.
Publicado: (2026)
ECLIPTICA -- A Framework for Switchable LLM Alignment via CITA - Contrastive Instruction-Tuned Alignment
por: Wanaskar, Kapil, et al.
Publicado: (2026)
por: Wanaskar, Kapil, et al.
Publicado: (2026)
AlignMerge - Alignment-Preserving Large Language Model Merging via Fisher-Guided Geometric Constraints
por: Roy, Aniruddha, et al.
Publicado: (2025)
por: Roy, Aniruddha, et al.
Publicado: (2025)
How Culturally Aware are Vision-Language Models?
por: Burda-Lassen, Olena, et al.
Publicado: (2024)
por: Burda-Lassen, Olena, et al.
Publicado: (2024)
Are Language Models Sensitive to Morally Irrelevant Distractors?
por: Shaw, Andrew, et al.
Publicado: (2026)
por: Shaw, Andrew, et al.
Publicado: (2026)
IMRNNs: An Efficient Method for Interpretable Dense Retrieval via Embedding Modulation
por: Saxena, Yash, et al.
Publicado: (2026)
por: Saxena, Yash, et al.
Publicado: (2026)
Context Matters: Auditing Gender Bias in T2I Generation through Risk-Tiered Use-Case Profiles
por: Luna, Jose, et al.
Publicado: (2026)
por: Luna, Jose, et al.
Publicado: (2026)
Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability
por: Raimondi, Bianca, et al.
Publicado: (2025)
por: Raimondi, Bianca, et al.
Publicado: (2025)
Neurosymbolic Retrievers for Retrieval-augmented Generation
por: Saxena, Yash, et al.
Publicado: (2026)
por: Saxena, Yash, et al.
Publicado: (2026)
Exploring the Frontier of Vision-Language Models: A Survey of Current Methodologies and Future Directions
por: Ghosh, Akash, et al.
Publicado: (2024)
por: Ghosh, Akash, et al.
Publicado: (2024)
A Comprehensive Survey of Accelerated Generation Techniques in Large Language Models
por: Khoshnoodi, Mahsa, et al.
Publicado: (2024)
por: Khoshnoodi, Mahsa, et al.
Publicado: (2024)
Assessing LLM Reliability on Temporally Recent Open-Domain Questions
por: Krishnappa, Pushwitha, et al.
Publicado: (2026)
por: Krishnappa, Pushwitha, et al.
Publicado: (2026)
Multilingual State Space Models for Structured Question Answering in Indic Languages
por: Vats, Arpita, et al.
Publicado: (2025)
por: Vats, Arpita, et al.
Publicado: (2025)
MOD-X: A Modular Open Decentralized eXchange Framework proposal for Heterogeneous Interoperable Artificial Intelligence Agents
por: Ioannides, Georgios, et al.
Publicado: (2025)
por: Ioannides, Georgios, et al.
Publicado: (2025)
Ejemplares similares
-
Unboxing Occupational Bias: Grounded Debiasing of LLMs with U.S. Labor Data
por: Gorti, Atmika, et al.
Publicado: (2024) -
Mental Health Equity in LLMs: Leveraging Multi-Hop Question Answering to Detect Amplified and Silenced Perspectives
por: Haider, Batool, et al.
Publicado: (2025) -
The Reasoning Trap -- Logical Reasoning as a Mechanistic Pathway to Situational Awareness
por: Sahoo, Subramanyam, et al.
Publicado: (2026) -
Born With a Silver Spoon? Investigating Socioeconomic Bias in Large Language Models
por: Singh, Smriti, et al.
Publicado: (2024) -
Dial E for Ethical Enforcement: institutional VETO power as a governance primitive
por: Sahoo, Subramanyam, et al.
Publicado: (2026)