Saved in:
| Main Authors: | Aghaebe, Favour Yahdii, Apekey, Tanefa, Williams, Elizabeth, Moosavi, Nafise Sadat |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2511.06000 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Faithful Summarisation under Disagreement via Belief-Level Aggregation
by: Aghaebe, Favour Yahdii, et al.
Published: (2026)
by: Aghaebe, Favour Yahdii, et al.
Published: (2026)
More or Less Wrong: A Benchmark for Directional Bias in LLM Comparative Reasoning
by: Shafiei, Mohammadamin, et al.
Published: (2025)
by: Shafiei, Mohammadamin, et al.
Published: (2025)
Exploring the Influence of Label Aggregation on Minority Voices: Implications for Dataset Bias and Model Training
by: Pandya, Mugdha, et al.
Published: (2024)
by: Pandya, Mugdha, et al.
Published: (2024)
LLMs as Narcissistic Evaluators: When Ego Inflates Evaluation Scores
by: Liu, Yiqi, et al.
Published: (2023)
by: Liu, Yiqi, et al.
Published: (2023)
Rolling the DICE on Idiomaticity: How LLMs Fail to Grasp Context
by: Mi, Maggie, et al.
Published: (2024)
by: Mi, Maggie, et al.
Published: (2024)
How to Leverage Digit Embeddings to Represent Numbers?
by: Sivakumar, Jasivan Alex, et al.
Published: (2024)
by: Sivakumar, Jasivan Alex, et al.
Published: (2024)
Initialisation Determines the Basin: Efficient Codebook Optimisation for Extreme LLM Quantization
by: Kennedy, Ian W., et al.
Published: (2026)
by: Kennedy, Ian W., et al.
Published: (2026)
MultiHoax: A Dataset of Multi-hop False-Premise Questions
by: Shafiei, Mohammadamin, et al.
Published: (2025)
by: Shafiei, Mohammadamin, et al.
Published: (2025)
From Input Perception to Predictive Insight: Modeling Model Blind Spots Before They Become Errors
by: Mi, Maggie, et al.
Published: (2025)
by: Mi, Maggie, et al.
Published: (2025)
Deconstructing Attention: Investigating Design Principles for Effective Language Modeling
by: Xue, Huiyin, et al.
Published: (2025)
by: Xue, Huiyin, et al.
Published: (2025)
Property Classification of Vacation Rental Properties during Covid-19
by: Aghaebe, Favour Yahdii, et al.
Published: (2025)
by: Aghaebe, Favour Yahdii, et al.
Published: (2025)
Decoding News Narratives: A Critical Analysis of Large Language Models in Framing Detection
by: Pastorino, Valeria, et al.
Published: (2024)
by: Pastorino, Valeria, et al.
Published: (2024)
No Shortcuts to Culture: Indonesian Multi-hop Question Answering for Complex Cultural Understanding
by: Permadi, Vynska Amalia, et al.
Published: (2026)
by: Permadi, Vynska Amalia, et al.
Published: (2026)
RIGOURATE: Quantifying Scientific Exaggeration with Evidence-Aligned Claim Evaluation
by: James, Joseph, et al.
Published: (2026)
by: James, Joseph, et al.
Published: (2026)
Beyond Hate Speech: NLP's Challenges and Opportunities in Uncovering Dehumanizing Language
by: Saffari, Hamidreza, et al.
Published: (2024)
by: Saffari, Hamidreza, et al.
Published: (2024)
Exploring Gender Disparities in Automatic Speech Recognition Technology
by: ElGhazaly, Hend, et al.
Published: (2025)
by: ElGhazaly, Hend, et al.
Published: (2025)
Hidden Failures in Robustness: Why Supervised Uncertainty Quantification Needs Better Evaluation
by: Stacey, Joe, et al.
Published: (2026)
by: Stacey, Joe, et al.
Published: (2026)
ContrastScore: Towards Higher Quality, Less Biased, More Efficient Evaluation Metrics with Contrastive Evaluation
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
Assessing the Reliability of LLMs Annotations in the Context of Demographic Bias and Model Explanation
by: Mohammadi, Hadi, et al.
Published: (2025)
by: Mohammadi, Hadi, et al.
Published: (2025)
Co‐Designing Recipe Resources to Support Healthy Eating in African‐Caribbeans in the United Kingdom: An Academic and Community Partnership Approach
by: Tanefa A. Apekey, et al.
Published: (2024)
by: Tanefa A. Apekey, et al.
Published: (2024)
Elucidating Mechanisms of Demographic Bias in LLMs for Healthcare
by: Ahsan, Hiba, et al.
Published: (2025)
by: Ahsan, Hiba, et al.
Published: (2025)
I Am Aligned, But With Whom? MENA Values Benchmark for Evaluating Cultural Alignment and Multilingual Bias in LLMs
by: Zahraei, Pardis Sadat, et al.
Published: (2025)
by: Zahraei, Pardis Sadat, et al.
Published: (2025)
The Media Bias Taxonomy: A Systematic Literature Review on the Forms and Automated Detection of Media Bias
by: Spinde, Timo, et al.
Published: (2023)
by: Spinde, Timo, et al.
Published: (2023)
Beyond Marginal Distributions: A Framework to Evaluate the Representativeness of Demographic-Aligned LLMs
by: Williams, Tristan, et al.
Published: (2026)
by: Williams, Tristan, et al.
Published: (2026)
Detecting Bias and Enhancing Diagnostic Accuracy in Large Language Models for Healthcare
by: Zahraei, Pardis Sadat, et al.
Published: (2024)
by: Zahraei, Pardis Sadat, et al.
Published: (2024)
Evaluating LLMs for Demographic-Targeted Social Bias Detection: A Comprehensive Benchmark Study
by: Majumdar, Ayan, et al.
Published: (2025)
by: Majumdar, Ayan, et al.
Published: (2025)
Translate With Care: Addressing Gender Bias, Neutrality, and Reasoning in Large Language Model Translations
by: Zahraei, Pardis Sadat, et al.
Published: (2025)
by: Zahraei, Pardis Sadat, et al.
Published: (2025)
Seeing Through AI's Lens: Enhancing Human Skepticism Towards LLM-Generated Fake News
by: Ayoobi, Navid, et al.
Published: (2024)
by: Ayoobi, Navid, et al.
Published: (2024)
Llama See, Llama Do: A Mechanistic Perspective on Contextual Entrainment and Distraction in LLMs
by: Niu, Jingcheng, et al.
Published: (2025)
by: Niu, Jingcheng, et al.
Published: (2025)
Sometimes the Model doth Preach: Quantifying Religious Bias in Open LLMs through Demographic Analysis in Asian Nations
by: Shankar, Hari, et al.
Published: (2025)
by: Shankar, Hari, et al.
Published: (2025)
Do Prevalent Bias Metrics Capture Allocational Harms from LLMs?
by: Cyberey, Hannah, et al.
Published: (2024)
by: Cyberey, Hannah, et al.
Published: (2024)
Different Bias Under Different Criteria: Assessing Bias in LLMs with a Fact-Based Approach
by: Ko, Changgeon, et al.
Published: (2024)
by: Ko, Changgeon, et al.
Published: (2024)
Bias in the Ear of the Listener: Assessing Sensitivity in Audio Language Models Across Linguistic, Demographic, and Positional Variations
by: Wei, Sheng-Lun, et al.
Published: (2026)
by: Wei, Sheng-Lun, et al.
Published: (2026)
LLMs Are Biased Towards Output Formats! Systematically Evaluating and Mitigating Output Format Bias of LLMs
by: Long, Do Xuan, et al.
Published: (2024)
by: Long, Do Xuan, et al.
Published: (2024)
ConsistencyAI: A Benchmark to Assess LLMs' Factual Consistency When Responding to Different Demographic Groups
by: Banyas, Peter, et al.
Published: (2025)
by: Banyas, Peter, et al.
Published: (2025)
Transforming Science with Large Language Models: A Survey on AI-assisted Scientific Discovery, Experimentation, Content Generation, and Evaluation
by: Eger, Steffen, et al.
Published: (2025)
by: Eger, Steffen, et al.
Published: (2025)
Measuring Mechanistic Independence: Can Bias Be Removed Without Erasing Demographics?
by: Shan, Zhengyang, et al.
Published: (2025)
by: Shan, Zhengyang, et al.
Published: (2025)
Do the Right Thing, Just Debias! Multi-Category Bias Mitigation Using LLMs
by: Roy, Amartya, et al.
Published: (2024)
by: Roy, Amartya, et al.
Published: (2024)
Demographic and Linguistic Bias Evaluation in Omnimodal Language Models
by: Elobaid, Alaa
Published: (2026)
by: Elobaid, Alaa
Published: (2026)
Do Large Language Models Reflect Demographic Pluralism in Safety?
by: Naseem, Usman, et al.
Published: (2026)
by: Naseem, Usman, et al.
Published: (2026)
Similar Items
-
Faithful Summarisation under Disagreement via Belief-Level Aggregation
by: Aghaebe, Favour Yahdii, et al.
Published: (2026) -
More or Less Wrong: A Benchmark for Directional Bias in LLM Comparative Reasoning
by: Shafiei, Mohammadamin, et al.
Published: (2025) -
Exploring the Influence of Label Aggregation on Minority Voices: Implications for Dataset Bias and Model Training
by: Pandya, Mugdha, et al.
Published: (2024) -
LLMs as Narcissistic Evaluators: When Ego Inflates Evaluation Scores
by: Liu, Yiqi, et al.
Published: (2023) -
Rolling the DICE on Idiomaticity: How LLMs Fail to Grasp Context
by: Mi, Maggie, et al.
Published: (2024)