Robust Bias Detection in MLMs and its Application to Human Trait Ratings
Fuente:
arXiv
Saved in:
| Main Authors: | Shrestha, Ingroj, Tay, Louis, Srinivasan, Padmini |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLM Bias Detection and Mitigation through the Lens of Desired Distributions
by: Shrestha, Ingroj, et al.
Published: (2025)
by: Shrestha, Ingroj, et al.
Published: (2025)
Sudden Drops in the Loss: Syntax Acquisition, Phase Transitions, and Simplicity Bias in MLMs
by: Chen, Angelica, et al.
Published: (2023)
by: Chen, Angelica, et al.
Published: (2023)
SLIM-LLMs: Modeling of Style-Sensory Language RelationshipsThrough Low-Dimensional Representations
by: Khalid, Osama, et al.
Published: (2025)
by: Khalid, Osama, et al.
Published: (2025)
Are MLMs Trapped in the Visual Room?
by: Zhang, Yazhou, et al.
Published: (2025)
by: Zhang, Yazhou, et al.
Published: (2025)
Efficient Fairness Testing in Large Language Models: Prioritizing Metamorphic Relations for Bias Detection
by: Giramata, Suavis, et al.
Published: (2025)
by: Giramata, Suavis, et al.
Published: (2025)
Inclusivity in Large Language Models: Personality Traits and Gender Bias in Scientific Abstracts
by: Pervez, Naseela, et al.
Published: (2024)
by: Pervez, Naseela, et al.
Published: (2024)
Incorporating Human Explanations for Robust Hate Speech Detection
by: Chen, Jennifer L., et al.
Published: (2024)
by: Chen, Jennifer L., et al.
Published: (2024)
Evaluating Language Model Character Traits
by: Ward, Francis Rhys, et al.
Published: (2024)
by: Ward, Francis Rhys, et al.
Published: (2024)
Metamorphic Testing for Fairness Evaluation in Large Language Models: Identifying Intersectional Bias in LLaMA and GPT
by: Reddy, Harishwar, et al.
Published: (2025)
by: Reddy, Harishwar, et al.
Published: (2025)
Decoding News Bias: Multi Bias Detection in News Articles
by: Shah, Bhushan Santosh, et al.
Published: (2025)
by: Shah, Bhushan Santosh, et al.
Published: (2025)
RuBia: A Russian Language Bias Detection Dataset
by: Grigoreva, Veronika, et al.
Published: (2024)
by: Grigoreva, Veronika, et al.
Published: (2024)
The Impact of Disability Disclosure on Fairness and Bias in LLM-Driven Candidate Selection
by: Kamruzzaman, Mahammed, et al.
Published: (2025)
by: Kamruzzaman, Mahammed, et al.
Published: (2025)
BiasAlert: A Plug-and-play Tool for Social Bias Detection in LLMs
by: Fan, Zhiting, et al.
Published: (2024)
by: Fan, Zhiting, et al.
Published: (2024)
On the Relationship between Truth and Political Bias in Language Models
by: Fulay, Suyash, et al.
Published: (2024)
by: Fulay, Suyash, et al.
Published: (2024)
Robust Bias Evaluation with FilBBQ: A Filipino Bias Benchmark for Question-Answering Language Models
by: Gamboa, Lance Calvin Lim, et al.
Published: (2026)
by: Gamboa, Lance Calvin Lim, et al.
Published: (2026)
BiasGuard: A Reasoning-enhanced Bias Detection Tool For Large Language Models
by: Fan, Zhiting, et al.
Published: (2025)
by: Fan, Zhiting, et al.
Published: (2025)
Trait-Aware Policy Optimization for Autoregressive Multi-Trait Essay Scoring
by: Wang, Zhengyang, et al.
Published: (2026)
by: Wang, Zhengyang, et al.
Published: (2026)
Bias Attribution in Filipino Language Models: Extending a Bias Interpretability Metric for Application on Agglutinative Languages
by: Gamboa, Lance Calvin Lim, et al.
Published: (2025)
by: Gamboa, Lance Calvin Lim, et al.
Published: (2025)
Improved Models for Media Bias Detection and Subcategorization
by: Menzner, Tim, et al.
Published: (2024)
by: Menzner, Tim, et al.
Published: (2024)
A Comparative Study of Large Language Models and Human Personality Traits
by: Jiaqi, Wang, et al.
Published: (2025)
by: Jiaqi, Wang, et al.
Published: (2025)
To Bias or Not to Bias: Detecting bias in News with bias-detector
by: Ghosh, Himel, et al.
Published: (2025)
by: Ghosh, Himel, et al.
Published: (2025)
BiasLab: Toward Explainable Political Bias Detection with Dual-Axis Annotations and Rationale Indicators
by: Solaiman, Kma
Published: (2025)
by: Solaiman, Kma
Published: (2025)
The Media Bias Taxonomy: A Systematic Literature Review on the Forms and Automated Detection of Media Bias
by: Spinde, Timo, et al.
Published: (2023)
by: Spinde, Timo, et al.
Published: (2023)
On Bias and Fairness in NLP: Investigating the Impact of Bias and Debiasing in Language Models on the Fairness of Toxicity Detection
by: Elsafoury, Fatma, et al.
Published: (2023)
by: Elsafoury, Fatma, et al.
Published: (2023)
Quantifying and Predicting Disagreement in Graded Human Ratings
by: Zhang, Leixin, et al.
Published: (2026)
by: Zhang, Leixin, et al.
Published: (2026)
Prompting Techniques for Reducing Social Bias in LLMs through System 1 and System 2 Cognitive Processes
by: Kamruzzaman, Mahammed, et al.
Published: (2024)
by: Kamruzzaman, Mahammed, et al.
Published: (2024)
BanglaIPA: Towards Robust Text-to-IPA Transcription with Contextual Rewriting in Bengali
by: Hasan, Jakir, et al.
Published: (2026)
by: Hasan, Jakir, et al.
Published: (2026)
From Calibration to Collaboration: LLM Uncertainty Quantification Should Be More Human-Centered
by: Devic, Siddartha, et al.
Published: (2025)
by: Devic, Siddartha, et al.
Published: (2025)
Blind to the Human Touch: Overlap Bias in LLM-Based Summary Evaluation
by: Fang, Jiangnan, et al.
Published: (2026)
by: Fang, Jiangnan, et al.
Published: (2026)
IndiVec: An Exploration of Leveraging Large Language Models for Media Bias Detection with Fine-Grained Bias Indicators
by: Lin, Luyang, et al.
Published: (2024)
by: Lin, Luyang, et al.
Published: (2024)
A Comparative Study of Light-weight Language Models for PII Masking and their Deployment for Real Conversational Texts
by: Acharya, Prabigya, et al.
Published: (2025)
by: Acharya, Prabigya, et al.
Published: (2025)
Accuracy and Political Bias of News Source Credibility Ratings by Large Language Models
by: Yang, Kai-Cheng, et al.
Published: (2023)
by: Yang, Kai-Cheng, et al.
Published: (2023)
Toward Robust LLM-Based Judges: Taxonomic Bias Evaluation and Debiasing Optimization
by: Zhou, Hongli, et al.
Published: (2026)
by: Zhou, Hongli, et al.
Published: (2026)
The BIAS Detection Framework: Bias Detection in Word Embeddings and Language Models for European Languages
by: Puttick, Alexandre, et al.
Published: (2024)
by: Puttick, Alexandre, et al.
Published: (2024)
Detecting Gender Bias in Course Evaluations
by: Lindau, Sarah, et al.
Published: (2024)
by: Lindau, Sarah, et al.
Published: (2024)
C3PA: An Open Dataset of Expert-Annotated and Regulation-Aware Privacy Policies to Enable Scalable Regulatory Compliance Audits
by: Musa, Maaz Bin, et al.
Published: (2024)
by: Musa, Maaz Bin, et al.
Published: (2024)
Detection, Classification, and Mitigation of Gender Bias in Large Language Models
by: Cheng, Xiaoqing, et al.
Published: (2025)
by: Cheng, Xiaoqing, et al.
Published: (2025)
BiaSWE: An Expert Annotated Dataset for Misogyny Detection in Swedish
by: Kukk, Kätriin, et al.
Published: (2025)
by: Kukk, Kätriin, et al.
Published: (2025)
Augmenting Bias Detection in LLMs Using Topological Data Analysis
by: Varadarajan, Keshav, et al.
Published: (2025)
by: Varadarajan, Keshav, et al.
Published: (2025)
Investigating Subtler Biases in LLMs: Ageism, Beauty, Institutional, and Nationality Bias in Generative Models
by: Kamruzzaman, Mahammed, et al.
Published: (2023)
by: Kamruzzaman, Mahammed, et al.
Published: (2023)
Similar Items
-
LLM Bias Detection and Mitigation through the Lens of Desired Distributions
by: Shrestha, Ingroj, et al.
Published: (2025) -
Sudden Drops in the Loss: Syntax Acquisition, Phase Transitions, and Simplicity Bias in MLMs
by: Chen, Angelica, et al.
Published: (2023) -
SLIM-LLMs: Modeling of Style-Sensory Language RelationshipsThrough Low-Dimensional Representations
by: Khalid, Osama, et al.
Published: (2025) -
Are MLMs Trapped in the Visual Room?
by: Zhang, Yazhou, et al.
Published: (2025) -
Efficient Fairness Testing in Large Language Models: Prioritizing Metamorphic Relations for Bias Detection
by: Giramata, Suavis, et al.
Published: (2025)