What's Not Said Still Hurts: A Description-Based Evaluation Framework for Measuring Social Bias in LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pan, Jinhao, Raj, Chahat, Yao, Ziyu, Zhu, Ziwei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Bias Association Discovery Framework for Open-Ended LLM Generations
von: Pan, Jinhao, et al.
Veröffentlicht: (2025)
von: Pan, Jinhao, et al.
Veröffentlicht: (2025)
Talent or Luck? Evaluating Attribution Bias in Large Language Models
von: Raj, Chahat, et al.
Veröffentlicht: (2025)
von: Raj, Chahat, et al.
Veröffentlicht: (2025)
Breaking Bias, Building Bridges: Evaluation and Mitigation of Social Biases in LLMs via Contact Hypothesis
von: Raj, Chahat, et al.
Veröffentlicht: (2024)
von: Raj, Chahat, et al.
Veröffentlicht: (2024)
VIGNETTE: Socially Grounded Bias Evaluation for Vision-Language Models
von: Raj, Chahat, et al.
Veröffentlicht: (2025)
von: Raj, Chahat, et al.
Veröffentlicht: (2025)
BiasDora: Exploring Hidden Biased Associations in Vision-Language Models
von: Raj, Chahat, et al.
Veröffentlicht: (2024)
von: Raj, Chahat, et al.
Veröffentlicht: (2024)
Purdah and Patriarchy: Evaluating and Mitigating South Asian Biases in Open-Ended Multilingual LLM Generations
von: Rinki, Mamnuya, et al.
Veröffentlicht: (2025)
von: Rinki, Mamnuya, et al.
Veröffentlicht: (2025)
Measuring Political Bias in Large Language Models: What Is Said and How It Is Said
von: Bang, Yejin, et al.
Veröffentlicht: (2024)
von: Bang, Yejin, et al.
Veröffentlicht: (2024)
KnowBias: Mitigating Social Bias in LLMs via Know-Bias Neuron Enhancement
von: Pan, Jinhao, et al.
Veröffentlicht: (2026)
von: Pan, Jinhao, et al.
Veröffentlicht: (2026)
They Said Memes Were Harmless-We Found the Ones That Hurt: Decoding Jokes, Symbols, and Cultural References
von: Tripathi, Sahil, et al.
Veröffentlicht: (2026)
von: Tripathi, Sahil, et al.
Veröffentlicht: (2026)
Evaluating Social Bias in RAG Systems: When External Context Helps and Reasoning Hurts
von: Parihar, Shweta, et al.
Veröffentlicht: (2026)
von: Parihar, Shweta, et al.
Veröffentlicht: (2026)
Ask LLMs Directly, "What shapes your bias?": Measuring Social Bias in Large Language Models
von: Shin, Jisu, et al.
Veröffentlicht: (2024)
von: Shin, Jisu, et al.
Veröffentlicht: (2024)
Navigating the Shortcut Maze: A Comprehensive Analysis of Shortcut Learning in Text Classification by Language Models
von: Zhou, Yuqing, et al.
Veröffentlicht: (2024)
von: Zhou, Yuqing, et al.
Veröffentlicht: (2024)
When Alignment Hurts: Decoupling Representational Spaces in Multilingual Models
von: Elshabrawy, Ahmed, et al.
Veröffentlicht: (2025)
von: Elshabrawy, Ahmed, et al.
Veröffentlicht: (2025)
Trustworthy Social Bias Measurement
von: Bommasani, Rishi, et al.
Veröffentlicht: (2022)
von: Bommasani, Rishi, et al.
Veröffentlicht: (2022)
Obscured but Not Erased: Evaluating Nationality Bias in LLMs via Name-Based Bias Benchmarks
von: Pelosio, Giulio, et al.
Veröffentlicht: (2025)
von: Pelosio, Giulio, et al.
Veröffentlicht: (2025)
Probing Social Identity Bias in Chinese LLMs with Gendered Pronouns and Social Groups
von: Liu, Geng, et al.
Veröffentlicht: (2025)
von: Liu, Geng, et al.
Veröffentlicht: (2025)
How Does Differential Privacy Affect Social Bias in LLMs? A Systematic Evaluation
von: Tenorio, Eduardo, et al.
Veröffentlicht: (2026)
von: Tenorio, Eduardo, et al.
Veröffentlicht: (2026)
FairCoder: Evaluating Social Bias of LLMs in Code Generation
von: Du, Yongkang, et al.
Veröffentlicht: (2025)
von: Du, Yongkang, et al.
Veröffentlicht: (2025)
Dual-Metric Evaluation of Social Bias in Large Language Models: Evidence from an Underrepresented Nepali Cultural Context
von: Pandey, Ashish, et al.
Veröffentlicht: (2026)
von: Pandey, Ashish, et al.
Veröffentlicht: (2026)
A Scalable Entity-Based Framework for Auditing Bias in LLMs
von: Elbouanani, Akram, et al.
Veröffentlicht: (2026)
von: Elbouanani, Akram, et al.
Veröffentlicht: (2026)
BiasAlert: A Plug-and-play Tool for Social Bias Detection in LLMs
von: Fan, Zhiting, et al.
Veröffentlicht: (2024)
von: Fan, Zhiting, et al.
Veröffentlicht: (2024)
FM SO.P: A Progressive Task Mixture Framework with Automatic Evaluation for Cross-Domain SOP Understanding
von: Huang, Siyuan, et al.
Veröffentlicht: (2026)
von: Huang, Siyuan, et al.
Veröffentlicht: (2026)
Instruction-Tuning LLMs for Event Extraction with Annotation Guidelines
von: Srivastava, Saurabh, et al.
Veröffentlicht: (2025)
von: Srivastava, Saurabh, et al.
Veröffentlicht: (2025)
Large Language Models Still Exhibit Bias in Long Text
von: Jeung, Wonje, et al.
Veröffentlicht: (2024)
von: Jeung, Wonje, et al.
Veröffentlicht: (2024)
When Reasoning Hurts: Source-Aware Evaluation of Frontier LLMs for Clinical SOAP Note Generation
von: Faisal, Faizan
Veröffentlicht: (2026)
von: Faisal, Faizan
Veröffentlicht: (2026)
CORTEX: Collaborative LLM Agents for High-Stakes Alert Triage
von: Wei, Bowen, et al.
Veröffentlicht: (2025)
von: Wei, Bowen, et al.
Veröffentlicht: (2025)
Evaluating LLMs for Demographic-Targeted Social Bias Detection: A Comprehensive Benchmark Study
von: Majumdar, Ayan, et al.
Veröffentlicht: (2025)
von: Majumdar, Ayan, et al.
Veröffentlicht: (2025)
A Japanese Benchmark for Evaluating Social Bias in Reasoning Based on Attribution Theory
von: Shiotani, Taihei, et al.
Veröffentlicht: (2026)
von: Shiotani, Taihei, et al.
Veröffentlicht: (2026)
SFT Doesn't Always Hurt General Capabilities: Revisiting Domain-Specific Fine-Tuning in LLMs
von: Lin, Jiacheng, et al.
Veröffentlicht: (2025)
von: Lin, Jiacheng, et al.
Veröffentlicht: (2025)
DIF: A Framework for Benchmarking and Verifying Implicit Bias in LLMs
von: Yin, Lake, et al.
Veröffentlicht: (2025)
von: Yin, Lake, et al.
Veröffentlicht: (2025)
Is a Peeled Apple Still Red? Evaluating LLMs' Ability for Conceptual Combination with Property Type
von: Song, Seokwon, et al.
Veröffentlicht: (2025)
von: Song, Seokwon, et al.
Veröffentlicht: (2025)
A Dual-Layered Evaluation of Geopolitical and Cultural Bias in LLMs
von: Kim, Sean, et al.
Veröffentlicht: (2025)
von: Kim, Sean, et al.
Veröffentlicht: (2025)
IndiCASA: A Dataset and Bias Evaluation Framework in LLMs Using Contrastive Embedding Similarity in the Indian Context
von: S, Santhosh G, et al.
Veröffentlicht: (2025)
von: S, Santhosh G, et al.
Veröffentlicht: (2025)
Evaluating Gender Bias of LLMs in Making Morality Judgements
von: Bajaj, Divij, et al.
Veröffentlicht: (2024)
von: Bajaj, Divij, et al.
Veröffentlicht: (2024)
Climbing the Ladder of Reasoning: What LLMs Can-and Still Can't-Solve after SFT?
von: Sun, Yiyou, et al.
Veröffentlicht: (2025)
von: Sun, Yiyou, et al.
Veröffentlicht: (2025)
Relative Bias: A Comparative Framework for Quantifying Bias in LLMs
von: Arbabi, Alireza, et al.
Veröffentlicht: (2025)
von: Arbabi, Alireza, et al.
Veröffentlicht: (2025)
Large Language Models Are Still Misled by Simple Bias Ensembles
von: Sun, Zhouhao, et al.
Veröffentlicht: (2025)
von: Sun, Zhouhao, et al.
Veröffentlicht: (2025)
Benford's Curse: Tracing Digit Bias to Numerical Hallucination in LLMs
von: Shao, Jiandong, et al.
Veröffentlicht: (2025)
von: Shao, Jiandong, et al.
Veröffentlicht: (2025)
User-Assistant Bias in LLMs
von: Pan, Xu, et al.
Veröffentlicht: (2025)
von: Pan, Xu, et al.
Veröffentlicht: (2025)
Evaluating the Bias in LLMs for Surveying Opinion and Decision Making in Healthcare
von: Khaokaew, Yonchanok, et al.
Veröffentlicht: (2025)
von: Khaokaew, Yonchanok, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Bias Association Discovery Framework for Open-Ended LLM Generations
von: Pan, Jinhao, et al.
Veröffentlicht: (2025) -
Talent or Luck? Evaluating Attribution Bias in Large Language Models
von: Raj, Chahat, et al.
Veröffentlicht: (2025) -
Breaking Bias, Building Bridges: Evaluation and Mitigation of Social Biases in LLMs via Contact Hypothesis
von: Raj, Chahat, et al.
Veröffentlicht: (2024) -
VIGNETTE: Socially Grounded Bias Evaluation for Vision-Language Models
von: Raj, Chahat, et al.
Veröffentlicht: (2025) -
BiasDora: Exploring Hidden Biased Associations in Vision-Language Models
von: Raj, Chahat, et al.
Veröffentlicht: (2024)