Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chandna, Bhavik, Bashir, Zubair, Sen, Procheta |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Counterfactual Explanation Framework for Retrieval Models
von: Chandna, Bhavik, et al.
Veröffentlicht: (2024)
von: Chandna, Bhavik, et al.
Veröffentlicht: (2024)
Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability
von: Raimondi, Bianca, et al.
Veröffentlicht: (2025)
von: Raimondi, Bianca, et al.
Veröffentlicht: (2025)
Mechanics of Bias and Reasoning: Interpreting the Impact of Chain-of-Thought Prompting on Gender Bias in LLMs
von: Pearman, Edie, et al.
Veröffentlicht: (2026)
von: Pearman, Edie, et al.
Veröffentlicht: (2026)
Mechanistic Interpretability of Cognitive Complexity in LLMs via Linear Probing using Bloom's Taxonomy
von: Raimondi, Bianca, et al.
Veröffentlicht: (2026)
von: Raimondi, Bianca, et al.
Veröffentlicht: (2026)
Mechanistic Interpretability Needs Philosophy
von: Williams, Iwan, et al.
Veröffentlicht: (2025)
von: Williams, Iwan, et al.
Veröffentlicht: (2025)
Evaluate Bias without Manual Test Sets: A Concept Representation Perspective for LLMs
von: Gao, Lang, et al.
Veröffentlicht: (2025)
von: Gao, Lang, et al.
Veröffentlicht: (2025)
Intrinsic Self-Correction in LLMs: Towards Explainable Prompting via Mechanistic Interpretability
von: Lee, Yu-Ting, et al.
Veröffentlicht: (2025)
von: Lee, Yu-Ting, et al.
Veröffentlicht: (2025)
Mechanistic Interpretability of Emotion Inference in Large Language Models
von: Tak, Ala N., et al.
Veröffentlicht: (2025)
von: Tak, Ala N., et al.
Veröffentlicht: (2025)
MIB: A Mechanistic Interpretability Benchmark
von: Mueller, Aaron, et al.
Veröffentlicht: (2025)
von: Mueller, Aaron, et al.
Veröffentlicht: (2025)
Towards Ethical Multi-Agent Systems of Large Language Models: A Mechanistic Interpretability Perspective
von: Lee, Jae Hee, et al.
Veröffentlicht: (2025)
von: Lee, Jae Hee, et al.
Veröffentlicht: (2025)
Eye of Judgement: Dissecting the Evaluation of Russian-speaking LLMs with POLLUX
von: Martynov, Nikita, et al.
Veröffentlicht: (2025)
von: Martynov, Nikita, et al.
Veröffentlicht: (2025)
Dissecting Role Cognition in Medical LLMs via Neuronal Ablation
von: Liang, Xun, et al.
Veröffentlicht: (2025)
von: Liang, Xun, et al.
Veröffentlicht: (2025)
The Language Barrier: Dissecting Safety Challenges of LLMs in Multilingual Contexts
von: Shen, Lingfeng, et al.
Veröffentlicht: (2024)
von: Shen, Lingfeng, et al.
Veröffentlicht: (2024)
Toward a Theory of Generalizability in LLM Mechanistic Interpretability Research
von: Trott, Sean
Veröffentlicht: (2025)
von: Trott, Sean
Veröffentlicht: (2025)
ExtremeAIGC: Benchmarking LMM Vulnerability to AI-Generated Extremist Content
von: Chandna, Bhavik, et al.
Veröffentlicht: (2025)
von: Chandna, Bhavik, et al.
Veröffentlicht: (2025)
Implicit Bias in LLMs: A Survey
von: Lin, Xinru, et al.
Veröffentlicht: (2025)
von: Lin, Xinru, et al.
Veröffentlicht: (2025)
Shifting Perspectives: Steering Vectors for Robust Bias Mitigation in LLMs
von: Siddique, Zara, et al.
Veröffentlicht: (2025)
von: Siddique, Zara, et al.
Veröffentlicht: (2025)
Mechanistic Interpretability of Socio-Political Frames in Language Models
von: Asghari, Hadi, et al.
Veröffentlicht: (2025)
von: Asghari, Hadi, et al.
Veröffentlicht: (2025)
Capturing Bias Diversity in LLMs
von: Gosavi, Purva Prasad, et al.
Veröffentlicht: (2024)
von: Gosavi, Purva Prasad, et al.
Veröffentlicht: (2024)
Mechanistic Interpretability of GPT-2: Lexical and Contextual Layers in Sentiment Analysis
von: Hatua, Amartya
Veröffentlicht: (2025)
von: Hatua, Amartya
Veröffentlicht: (2025)
Towards Understanding and Improving Refusal in Compressed Models via Mechanistic Interpretability
von: Chhabra, Vishnu Kabir, et al.
Veröffentlicht: (2025)
von: Chhabra, Vishnu Kabir, et al.
Veröffentlicht: (2025)
How Emotion Shapes the Behavior of LLMs and Agents: A Mechanistic Study
von: Sun, Moran, et al.
Veröffentlicht: (2026)
von: Sun, Moran, et al.
Veröffentlicht: (2026)
Mechanistic Interpretability as Statistical Estimation: A Variance Analysis
von: Méloux, Maxime, et al.
Veröffentlicht: (2025)
von: Méloux, Maxime, et al.
Veröffentlicht: (2025)
Cognitive Bias in Decision-Making with LLMs
von: Echterhoff, Jessica, et al.
Veröffentlicht: (2024)
von: Echterhoff, Jessica, et al.
Veröffentlicht: (2024)
Think Inside the JSON: Reinforcement Strategy for Strict LLM Schema Adherence
von: Agarwal, Bhavik, et al.
Veröffentlicht: (2025)
von: Agarwal, Bhavik, et al.
Veröffentlicht: (2025)
What Defines Good Reasoning in LLMs? Dissecting Reasoning Steps with Multi-Aspect Evaluation
von: Do, Heejin, et al.
Veröffentlicht: (2025)
von: Do, Heejin, et al.
Veröffentlicht: (2025)
Safe in the Future, Dangerous in the Past: Dissecting Temporal and Linguistic Vulnerabilities in LLMs
von: Said, Muhammad Abdullahi, et al.
Veröffentlicht: (2025)
von: Said, Muhammad Abdullahi, et al.
Veröffentlicht: (2025)
One Bias After Another: Mechanistic Reward Shaping and Persistent Biases in Language Reward Models
von: Fein, Daniel, et al.
Veröffentlicht: (2026)
von: Fein, Daniel, et al.
Veröffentlicht: (2026)
XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content
von: Abishethvarman, Vadivel, et al.
Veröffentlicht: (2025)
von: Abishethvarman, Vadivel, et al.
Veröffentlicht: (2025)
LLMs in Interpreting Legal Documents
von: Corbo, Simone
Veröffentlicht: (2025)
von: Corbo, Simone
Veröffentlicht: (2025)
Detection and Mitigation of Hallucination in Large Reasoning Models: A Mechanistic Perspective
von: Sun, Zhongxiang, et al.
Veröffentlicht: (2025)
von: Sun, Zhongxiang, et al.
Veröffentlicht: (2025)
IndoBias: A Dual Track Culturally Grounded Benchmark for LLMs Bias Evaluation in Indonesian Languages
von: Hanif, Ikhlasul Akmal, et al.
Veröffentlicht: (2026)
von: Hanif, Ikhlasul Akmal, et al.
Veröffentlicht: (2026)
Voice, Bias, and Coreference: An Interpretability Study of Gender in Speech Translation
von: Conti, Lina, et al.
Veröffentlicht: (2025)
von: Conti, Lina, et al.
Veröffentlicht: (2025)
A Scalable Entity-Based Framework for Auditing Bias in LLMs
von: Elbouanani, Akram, et al.
Veröffentlicht: (2026)
von: Elbouanani, Akram, et al.
Veröffentlicht: (2026)
HyperDAS: Towards Automating Mechanistic Interpretability with Hypernetworks
von: Sun, Jiuding, et al.
Veröffentlicht: (2025)
von: Sun, Jiuding, et al.
Veröffentlicht: (2025)
Mechanistic Interpretability of GPT-like Models on Summarization Tasks
von: Mishra, Anurag
Veröffentlicht: (2025)
von: Mishra, Anurag
Veröffentlicht: (2025)
Binary Autoencoder for Mechanistic Interpretability of Large Language Models
von: Cho, Hakaze, et al.
Veröffentlicht: (2025)
von: Cho, Hakaze, et al.
Veröffentlicht: (2025)
Everything, Everywhere, All at Once: Is Mechanistic Interpretability Identifiable?
von: Méloux, Maxime, et al.
Veröffentlicht: (2025)
von: Méloux, Maxime, et al.
Veröffentlicht: (2025)
Dissecting Logical Reasoning in LLMs: A Fine-Grained Evaluation and Supervision Study
von: Zhou, Yujun, et al.
Veröffentlicht: (2025)
von: Zhou, Yujun, et al.
Veröffentlicht: (2025)
Mechanistic Steering of LLMs Reveals Layer-wise Feature Vulnerabilities in Adversarial Settings
von: Das, Nilanjana, et al.
Veröffentlicht: (2026)
von: Das, Nilanjana, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
A Counterfactual Explanation Framework for Retrieval Models
von: Chandna, Bhavik, et al.
Veröffentlicht: (2024) -
Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability
von: Raimondi, Bianca, et al.
Veröffentlicht: (2025) -
Mechanics of Bias and Reasoning: Interpreting the Impact of Chain-of-Thought Prompting on Gender Bias in LLMs
von: Pearman, Edie, et al.
Veröffentlicht: (2026) -
Mechanistic Interpretability of Cognitive Complexity in LLMs via Linear Probing using Bloom's Taxonomy
von: Raimondi, Bianca, et al.
Veröffentlicht: (2026) -
Mechanistic Interpretability Needs Philosophy
von: Williams, Iwan, et al.
Veröffentlicht: (2025)