Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective
Fuente:
arXiv
Salvato in:
| Autori principali: | Chandna, Bhavik, Bashir, Zubair, Sen, Procheta |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A Counterfactual Explanation Framework for Retrieval Models
di: Chandna, Bhavik, et al.
Pubblicazione: (2024)
di: Chandna, Bhavik, et al.
Pubblicazione: (2024)
Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability
di: Raimondi, Bianca, et al.
Pubblicazione: (2025)
di: Raimondi, Bianca, et al.
Pubblicazione: (2025)
Mechanics of Bias and Reasoning: Interpreting the Impact of Chain-of-Thought Prompting on Gender Bias in LLMs
di: Pearman, Edie, et al.
Pubblicazione: (2026)
di: Pearman, Edie, et al.
Pubblicazione: (2026)
Mechanistic Interpretability of Cognitive Complexity in LLMs via Linear Probing using Bloom's Taxonomy
di: Raimondi, Bianca, et al.
Pubblicazione: (2026)
di: Raimondi, Bianca, et al.
Pubblicazione: (2026)
Mechanistic Interpretability Needs Philosophy
di: Williams, Iwan, et al.
Pubblicazione: (2025)
di: Williams, Iwan, et al.
Pubblicazione: (2025)
Evaluate Bias without Manual Test Sets: A Concept Representation Perspective for LLMs
di: Gao, Lang, et al.
Pubblicazione: (2025)
di: Gao, Lang, et al.
Pubblicazione: (2025)
Intrinsic Self-Correction in LLMs: Towards Explainable Prompting via Mechanistic Interpretability
di: Lee, Yu-Ting, et al.
Pubblicazione: (2025)
di: Lee, Yu-Ting, et al.
Pubblicazione: (2025)
Mechanistic Interpretability of Emotion Inference in Large Language Models
di: Tak, Ala N., et al.
Pubblicazione: (2025)
di: Tak, Ala N., et al.
Pubblicazione: (2025)
MIB: A Mechanistic Interpretability Benchmark
di: Mueller, Aaron, et al.
Pubblicazione: (2025)
di: Mueller, Aaron, et al.
Pubblicazione: (2025)
Towards Ethical Multi-Agent Systems of Large Language Models: A Mechanistic Interpretability Perspective
di: Lee, Jae Hee, et al.
Pubblicazione: (2025)
di: Lee, Jae Hee, et al.
Pubblicazione: (2025)
Eye of Judgement: Dissecting the Evaluation of Russian-speaking LLMs with POLLUX
di: Martynov, Nikita, et al.
Pubblicazione: (2025)
di: Martynov, Nikita, et al.
Pubblicazione: (2025)
Dissecting Role Cognition in Medical LLMs via Neuronal Ablation
di: Liang, Xun, et al.
Pubblicazione: (2025)
di: Liang, Xun, et al.
Pubblicazione: (2025)
The Language Barrier: Dissecting Safety Challenges of LLMs in Multilingual Contexts
di: Shen, Lingfeng, et al.
Pubblicazione: (2024)
di: Shen, Lingfeng, et al.
Pubblicazione: (2024)
Toward a Theory of Generalizability in LLM Mechanistic Interpretability Research
di: Trott, Sean
Pubblicazione: (2025)
di: Trott, Sean
Pubblicazione: (2025)
ExtremeAIGC: Benchmarking LMM Vulnerability to AI-Generated Extremist Content
di: Chandna, Bhavik, et al.
Pubblicazione: (2025)
di: Chandna, Bhavik, et al.
Pubblicazione: (2025)
Implicit Bias in LLMs: A Survey
di: Lin, Xinru, et al.
Pubblicazione: (2025)
di: Lin, Xinru, et al.
Pubblicazione: (2025)
Shifting Perspectives: Steering Vectors for Robust Bias Mitigation in LLMs
di: Siddique, Zara, et al.
Pubblicazione: (2025)
di: Siddique, Zara, et al.
Pubblicazione: (2025)
Mechanistic Interpretability of Socio-Political Frames in Language Models
di: Asghari, Hadi, et al.
Pubblicazione: (2025)
di: Asghari, Hadi, et al.
Pubblicazione: (2025)
Capturing Bias Diversity in LLMs
di: Gosavi, Purva Prasad, et al.
Pubblicazione: (2024)
di: Gosavi, Purva Prasad, et al.
Pubblicazione: (2024)
Mechanistic Interpretability of GPT-2: Lexical and Contextual Layers in Sentiment Analysis
di: Hatua, Amartya
Pubblicazione: (2025)
di: Hatua, Amartya
Pubblicazione: (2025)
Towards Understanding and Improving Refusal in Compressed Models via Mechanistic Interpretability
di: Chhabra, Vishnu Kabir, et al.
Pubblicazione: (2025)
di: Chhabra, Vishnu Kabir, et al.
Pubblicazione: (2025)
How Emotion Shapes the Behavior of LLMs and Agents: A Mechanistic Study
di: Sun, Moran, et al.
Pubblicazione: (2026)
di: Sun, Moran, et al.
Pubblicazione: (2026)
Mechanistic Interpretability as Statistical Estimation: A Variance Analysis
di: Méloux, Maxime, et al.
Pubblicazione: (2025)
di: Méloux, Maxime, et al.
Pubblicazione: (2025)
Cognitive Bias in Decision-Making with LLMs
di: Echterhoff, Jessica, et al.
Pubblicazione: (2024)
di: Echterhoff, Jessica, et al.
Pubblicazione: (2024)
Think Inside the JSON: Reinforcement Strategy for Strict LLM Schema Adherence
di: Agarwal, Bhavik, et al.
Pubblicazione: (2025)
di: Agarwal, Bhavik, et al.
Pubblicazione: (2025)
What Defines Good Reasoning in LLMs? Dissecting Reasoning Steps with Multi-Aspect Evaluation
di: Do, Heejin, et al.
Pubblicazione: (2025)
di: Do, Heejin, et al.
Pubblicazione: (2025)
Safe in the Future, Dangerous in the Past: Dissecting Temporal and Linguistic Vulnerabilities in LLMs
di: Said, Muhammad Abdullahi, et al.
Pubblicazione: (2025)
di: Said, Muhammad Abdullahi, et al.
Pubblicazione: (2025)
One Bias After Another: Mechanistic Reward Shaping and Persistent Biases in Language Reward Models
di: Fein, Daniel, et al.
Pubblicazione: (2026)
di: Fein, Daniel, et al.
Pubblicazione: (2026)
XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content
di: Abishethvarman, Vadivel, et al.
Pubblicazione: (2025)
di: Abishethvarman, Vadivel, et al.
Pubblicazione: (2025)
LLMs in Interpreting Legal Documents
di: Corbo, Simone
Pubblicazione: (2025)
di: Corbo, Simone
Pubblicazione: (2025)
Detection and Mitigation of Hallucination in Large Reasoning Models: A Mechanistic Perspective
di: Sun, Zhongxiang, et al.
Pubblicazione: (2025)
di: Sun, Zhongxiang, et al.
Pubblicazione: (2025)
IndoBias: A Dual Track Culturally Grounded Benchmark for LLMs Bias Evaluation in Indonesian Languages
di: Hanif, Ikhlasul Akmal, et al.
Pubblicazione: (2026)
di: Hanif, Ikhlasul Akmal, et al.
Pubblicazione: (2026)
Voice, Bias, and Coreference: An Interpretability Study of Gender in Speech Translation
di: Conti, Lina, et al.
Pubblicazione: (2025)
di: Conti, Lina, et al.
Pubblicazione: (2025)
A Scalable Entity-Based Framework for Auditing Bias in LLMs
di: Elbouanani, Akram, et al.
Pubblicazione: (2026)
di: Elbouanani, Akram, et al.
Pubblicazione: (2026)
HyperDAS: Towards Automating Mechanistic Interpretability with Hypernetworks
di: Sun, Jiuding, et al.
Pubblicazione: (2025)
di: Sun, Jiuding, et al.
Pubblicazione: (2025)
Mechanistic Interpretability of GPT-like Models on Summarization Tasks
di: Mishra, Anurag
Pubblicazione: (2025)
di: Mishra, Anurag
Pubblicazione: (2025)
Binary Autoencoder for Mechanistic Interpretability of Large Language Models
di: Cho, Hakaze, et al.
Pubblicazione: (2025)
di: Cho, Hakaze, et al.
Pubblicazione: (2025)
Everything, Everywhere, All at Once: Is Mechanistic Interpretability Identifiable?
di: Méloux, Maxime, et al.
Pubblicazione: (2025)
di: Méloux, Maxime, et al.
Pubblicazione: (2025)
Dissecting Logical Reasoning in LLMs: A Fine-Grained Evaluation and Supervision Study
di: Zhou, Yujun, et al.
Pubblicazione: (2025)
di: Zhou, Yujun, et al.
Pubblicazione: (2025)
Mechanistic Steering of LLMs Reveals Layer-wise Feature Vulnerabilities in Adversarial Settings
di: Das, Nilanjana, et al.
Pubblicazione: (2026)
di: Das, Nilanjana, et al.
Pubblicazione: (2026)
Documenti analoghi
-
A Counterfactual Explanation Framework for Retrieval Models
di: Chandna, Bhavik, et al.
Pubblicazione: (2024) -
Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability
di: Raimondi, Bianca, et al.
Pubblicazione: (2025) -
Mechanics of Bias and Reasoning: Interpreting the Impact of Chain-of-Thought Prompting on Gender Bias in LLMs
di: Pearman, Edie, et al.
Pubblicazione: (2026) -
Mechanistic Interpretability of Cognitive Complexity in LLMs via Linear Probing using Bloom's Taxonomy
di: Raimondi, Bianca, et al.
Pubblicazione: (2026) -
Mechanistic Interpretability Needs Philosophy
di: Williams, Iwan, et al.
Pubblicazione: (2025)