Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability
Fuente:
arXiv
Saved in:
| Main Authors: | Raimondi, Bianca, Dalbagno, Daniela, Gabbrielli, Maurizio |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mechanistic Interpretability of Cognitive Complexity in LLMs via Linear Probing using Bloom's Taxonomy
by: Raimondi, Bianca, et al.
Published: (2026)
by: Raimondi, Bianca, et al.
Published: (2026)
Exploiting Primacy Effect To Improve Large Language Models
by: Raimondi, Bianca, et al.
Published: (2025)
by: Raimondi, Bianca, et al.
Published: (2025)
Affordably Fine-tuned LLMs Provide Better Answers to Course-specific MCQs
by: Raimondi, Bianca, et al.
Published: (2025)
by: Raimondi, Bianca, et al.
Published: (2025)
Learning Factors in AI-Augmented Education: A Comparative Study of Middle and High School Students
by: Ebli, Gaia, et al.
Published: (2025)
by: Ebli, Gaia, et al.
Published: (2025)
From Reasoning to Code: GRPO Optimization for Underrepresented Languages
by: Pennino, Federico, et al.
Published: (2025)
by: Pennino, Federico, et al.
Published: (2025)
The CompMath-MCQ Dataset: Are LLMs Ready for Higher-Level Math?
by: Raimondi, Bianca, et al.
Published: (2026)
by: Raimondi, Bianca, et al.
Published: (2026)
Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective
by: Chandna, Bhavik, et al.
Published: (2025)
by: Chandna, Bhavik, et al.
Published: (2025)
MoSE: Hierarchical Self-Distillation Enhances Early Layer Embeddings
by: Gurioli, Andrea, et al.
Published: (2025)
by: Gurioli, Andrea, et al.
Published: (2025)
Political Bias in LLMs: Unaligned Moral Values in Agent-centric Simulations
by: Münker, Simon
Published: (2024)
by: Münker, Simon
Published: (2024)
Evaluation of Finetuned LLMs in AMR Parsing
by: Ho, Shu Han
Published: (2025)
by: Ho, Shu Han
Published: (2025)
Widespread Gender and Pronoun Bias in Moral Judgments Across LLMs
by: Fernandes, Gustavo Lúcius, et al.
Published: (2026)
by: Fernandes, Gustavo Lúcius, et al.
Published: (2026)
Mechanics of Bias and Reasoning: Interpreting the Impact of Chain-of-Thought Prompting on Gender Bias in LLMs
by: Pearman, Edie, et al.
Published: (2026)
by: Pearman, Edie, et al.
Published: (2026)
Common Sense vs. Morality: The Curious Case of Narrative Focus Bias in LLMs
by: Purkayastha, Saugata, et al.
Published: (2026)
by: Purkayastha, Saugata, et al.
Published: (2026)
High Accuracy, Less Talk (HALT): Reliable LLMs through Capability-Aligned Finetuning
by: Franzmeyer, Tim, et al.
Published: (2025)
by: Franzmeyer, Tim, et al.
Published: (2025)
Mechanistic Origin of Moral Indifference in Language Models
by: Li, Lingyu, et al.
Published: (2026)
by: Li, Lingyu, et al.
Published: (2026)
Mechanistic Interpretability Needs Philosophy
by: Williams, Iwan, et al.
Published: (2025)
by: Williams, Iwan, et al.
Published: (2025)
MoralBench: Moral Evaluation of LLMs
by: Ji, Jianchao, et al.
Published: (2024)
by: Ji, Jianchao, et al.
Published: (2024)
Cultural Bias in Large Language Models: Evaluating AI Agents through Moral Questionnaires
by: Münker, Simon
Published: (2025)
by: Münker, Simon
Published: (2025)
Improving Task Diversity in Label Efficient Supervised Finetuning of LLMs
by: Arabelly, Abhinav, et al.
Published: (2025)
by: Arabelly, Abhinav, et al.
Published: (2025)
Evaluating Moral Beliefs across LLMs through a Pluralistic Framework
by: Liu, Xuelin, et al.
Published: (2024)
by: Liu, Xuelin, et al.
Published: (2024)
Intrinsic Self-Correction in LLMs: Towards Explainable Prompting via Mechanistic Interpretability
by: Lee, Yu-Ting, et al.
Published: (2025)
by: Lee, Yu-Ting, et al.
Published: (2025)
The LLM Wears Prada: Analysing Gender Bias and Stereotypes through Online Shopping Data
by: Luca, Massimiliano, et al.
Published: (2025)
by: Luca, Massimiliano, et al.
Published: (2025)
Mechanistic Interpretability of Emotion Inference in Large Language Models
by: Tak, Ala N., et al.
Published: (2025)
by: Tak, Ala N., et al.
Published: (2025)
Detecting and Mitigating Bias in LLMs through Knowledge Graph-Augmented Training
by: Kumar, Rajeev, et al.
Published: (2025)
by: Kumar, Rajeev, et al.
Published: (2025)
Toward a Theory of Generalizability in LLM Mechanistic Interpretability Research
by: Trott, Sean
Published: (2025)
by: Trott, Sean
Published: (2025)
mdok-style at SemEval-2026 Task 10: Finetuning LLMs for Conspiracy Detection
by: Macko, Dominik
Published: (2026)
by: Macko, Dominik
Published: (2026)
Mechanistic Interpretability of Socio-Political Frames in Language Models
by: Asghari, Hadi, et al.
Published: (2025)
by: Asghari, Hadi, et al.
Published: (2025)
Capturing Bias Diversity in LLMs
by: Gosavi, Purva Prasad, et al.
Published: (2024)
by: Gosavi, Purva Prasad, et al.
Published: (2024)
MIB: A Mechanistic Interpretability Benchmark
by: Mueller, Aaron, et al.
Published: (2025)
by: Mueller, Aaron, et al.
Published: (2025)
Robo-Instruct: Simulator-Augmented Instruction Alignment For Finetuning Code LLMs
by: Hu, Zichao, et al.
Published: (2024)
by: Hu, Zichao, et al.
Published: (2024)
Mechanistic Interpretability of GPT-2: Lexical and Contextual Layers in Sentiment Analysis
by: Hatua, Amartya
Published: (2025)
by: Hatua, Amartya
Published: (2025)
Towards Understanding and Improving Refusal in Compressed Models via Mechanistic Interpretability
by: Chhabra, Vishnu Kabir, et al.
Published: (2025)
by: Chhabra, Vishnu Kabir, et al.
Published: (2025)
ProMoral-Bench: Evaluating Prompting Strategies for Moral Reasoning and Safety in LLMs
by: Thomas, Rohan Subramanian, et al.
Published: (2026)
by: Thomas, Rohan Subramanian, et al.
Published: (2026)
Moral Lenses, Political Coordinates: Towards Ideological Positioning of Morally Conditioned LLMs
by: Yuan, Chenchen, et al.
Published: (2026)
by: Yuan, Chenchen, et al.
Published: (2026)
Moral Mazes in the Era of LLMs
by: Nguyen, Dang, et al.
Published: (2026)
by: Nguyen, Dang, et al.
Published: (2026)
How to Alleviate Catastrophic Forgetting in LLMs Finetuning? Hierarchical Layer-Wise and Element-Wise Regularization
by: Song, Shezheng, et al.
Published: (2025)
by: Song, Shezheng, et al.
Published: (2025)
mdok-style at SemEval-2026 Task 9: Finetuning LLMs for Multilingual Polarization Detection
by: Macko, Dominik, et al.
Published: (2026)
by: Macko, Dominik, et al.
Published: (2026)
Exploring the psychology of LLMs' Moral and Legal Reasoning
by: Almeida, Guilherme F. C. F., et al.
Published: (2023)
by: Almeida, Guilherme F. C. F., et al.
Published: (2023)
Analyzing Finetuning Representation Shift for Multimodal LLMs Steering
by: Khayatan, Pegah, et al.
Published: (2025)
by: Khayatan, Pegah, et al.
Published: (2025)
Implicit Bias in LLMs: A Survey
by: Lin, Xinru, et al.
Published: (2025)
by: Lin, Xinru, et al.
Published: (2025)
Similar Items
-
Mechanistic Interpretability of Cognitive Complexity in LLMs via Linear Probing using Bloom's Taxonomy
by: Raimondi, Bianca, et al.
Published: (2026) -
Exploiting Primacy Effect To Improve Large Language Models
by: Raimondi, Bianca, et al.
Published: (2025) -
Affordably Fine-tuned LLMs Provide Better Answers to Course-specific MCQs
by: Raimondi, Bianca, et al.
Published: (2025) -
Learning Factors in AI-Augmented Education: A Comparative Study of Middle and High School Students
by: Ebli, Gaia, et al.
Published: (2025) -
From Reasoning to Code: GRPO Optimization for Underrepresented Languages
by: Pennino, Federico, et al.
Published: (2025)