BiasAlert: A Plug-and-play Tool for Social Bias Detection in LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fan, Zhiting, Chen, Ruizhe, Xu, Ruiling, Liu, Zuozhu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BiasGuard: A Reasoning-enhanced Bias Detection Tool For Large Language Models
von: Fan, Zhiting, et al.
Veröffentlicht: (2025)
von: Fan, Zhiting, et al.
Veröffentlicht: (2025)
FairMT-Bench: Benchmarking Fairness for Multi-turn Dialogue in Conversational LLMs
von: Fan, Zhiting, et al.
Veröffentlicht: (2024)
von: Fan, Zhiting, et al.
Veröffentlicht: (2024)
Large Language Model Bias Mitigation from the Perspective of Knowledge Editing
von: Chen, Ruizhe, et al.
Veröffentlicht: (2024)
von: Chen, Ruizhe, et al.
Veröffentlicht: (2024)
Identifying and Mitigating Social Bias Knowledge in Language Models
von: Chen, Ruizhe, et al.
Veröffentlicht: (2024)
von: Chen, Ruizhe, et al.
Veröffentlicht: (2024)
FairSteer: Inference Time Debiasing for LLMs with Dynamic Activation Steering
von: Li, Yichen, et al.
Veröffentlicht: (2025)
von: Li, Yichen, et al.
Veröffentlicht: (2025)
PAD: Personalized Alignment of LLMs at Decoding-Time
von: Chen, Ruizhe, et al.
Veröffentlicht: (2024)
von: Chen, Ruizhe, et al.
Veröffentlicht: (2024)
FAIntbench: A Holistic and Precise Benchmark for Bias Evaluation in Text-to-Image Models
von: Luo, Hanjun, et al.
Veröffentlicht: (2024)
von: Luo, Hanjun, et al.
Veröffentlicht: (2024)
Optimsyn: Influence-Guided Rubrics Optimization for Synthetic Data Generation
von: Fan, Zhiting, et al.
Veröffentlicht: (2026)
von: Fan, Zhiting, et al.
Veröffentlicht: (2026)
BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model Responses
von: Xu, Xin, et al.
Veröffentlicht: (2025)
von: Xu, Xin, et al.
Veröffentlicht: (2025)
User-Assistant Bias in LLMs
von: Pan, Xu, et al.
Veröffentlicht: (2025)
von: Pan, Xu, et al.
Veröffentlicht: (2025)
Are LLMs Rational Investors? A Study on Detecting and Reducing the Financial Bias in LLMs
von: Zhou, Yuhang, et al.
Veröffentlicht: (2024)
von: Zhou, Yuhang, et al.
Veröffentlicht: (2024)
Evaluating LLMs for Demographic-Targeted Social Bias Detection: A Comprehensive Benchmark Study
von: Majumdar, Ayan, et al.
Veröffentlicht: (2025)
von: Majumdar, Ayan, et al.
Veröffentlicht: (2025)
The Bias is in the Details: An Assessment of Cognitive Bias in LLMs
von: Knipper, R. Alexander, et al.
Veröffentlicht: (2025)
von: Knipper, R. Alexander, et al.
Veröffentlicht: (2025)
Persona-judge: Personalized Alignment of Large Language Models via Token-level Self-judgment
von: Zhang, Xiaotian, et al.
Veröffentlicht: (2025)
von: Zhang, Xiaotian, et al.
Veröffentlicht: (2025)
Do LLMs Know Tool Irrelevance? Demystifying Structural Alignment Bias in Tool Invocations
von: Liu, Yilong, et al.
Veröffentlicht: (2026)
von: Liu, Yilong, et al.
Veröffentlicht: (2026)
DIF: A Framework for Benchmarking and Verifying Implicit Bias in LLMs
von: Yin, Lake, et al.
Veröffentlicht: (2025)
von: Yin, Lake, et al.
Veröffentlicht: (2025)
Race, Ethnicity and Their Implication on Bias in Large Language Models
von: Hu, Shiyue, et al.
Veröffentlicht: (2026)
von: Hu, Shiyue, et al.
Veröffentlicht: (2026)
Probing Social Identity Bias in Chinese LLMs with Gendered Pronouns and Social Groups
von: Liu, Geng, et al.
Veröffentlicht: (2025)
von: Liu, Geng, et al.
Veröffentlicht: (2025)
BiasFilter: An Inference-Time Debiasing Framework for Large Language Models
von: Cheng, Xiaoqing, et al.
Veröffentlicht: (2025)
von: Cheng, Xiaoqing, et al.
Veröffentlicht: (2025)
Causally Testing Gender Bias in LLMs: A Case Study on Occupational Bias
von: Chen, Yuen, et al.
Veröffentlicht: (2022)
von: Chen, Yuen, et al.
Veröffentlicht: (2022)
RuBia: A Russian Language Bias Detection Dataset
von: Grigoreva, Veronika, et al.
Veröffentlicht: (2024)
von: Grigoreva, Veronika, et al.
Veröffentlicht: (2024)
Anchored Answers: Unravelling Positional Bias in GPT-2's Multiple-Choice Questions
von: Li, Ruizhe, et al.
Veröffentlicht: (2024)
von: Li, Ruizhe, et al.
Veröffentlicht: (2024)
Augmenting Bias Detection in LLMs Using Topological Data Analysis
von: Varadarajan, Keshav, et al.
Veröffentlicht: (2025)
von: Varadarajan, Keshav, et al.
Veröffentlicht: (2025)
Learnable Privacy Neurons Localization in Language Models
von: Chen, Ruizhe, et al.
Veröffentlicht: (2024)
von: Chen, Ruizhe, et al.
Veröffentlicht: (2024)
Decoding News Bias: Multi Bias Detection in News Articles
von: Shah, Bhushan Santosh, et al.
Veröffentlicht: (2025)
von: Shah, Bhushan Santosh, et al.
Veröffentlicht: (2025)
BiasIG: Benchmarking Multi-dimensional Social Biases in Text-to-Image Models
von: Luo, Hanjun, et al.
Veröffentlicht: (2026)
von: Luo, Hanjun, et al.
Veröffentlicht: (2026)
Relative Bias: A Comparative Framework for Quantifying Bias in LLMs
von: Arbabi, Alireza, et al.
Veröffentlicht: (2025)
von: Arbabi, Alireza, et al.
Veröffentlicht: (2025)
Disclosure and Mitigation of Gender Bias in LLMs
von: Dong, Xiangjue, et al.
Veröffentlicht: (2024)
von: Dong, Xiangjue, et al.
Veröffentlicht: (2024)
Cognitive Bias in Decision-Making with LLMs
von: Echterhoff, Jessica, et al.
Veröffentlicht: (2024)
von: Echterhoff, Jessica, et al.
Veröffentlicht: (2024)
FairCoder: Evaluating Social Bias of LLMs in Code Generation
von: Du, Yongkang, et al.
Veröffentlicht: (2025)
von: Du, Yongkang, et al.
Veröffentlicht: (2025)
Ladder: A Model-Agnostic Framework Boosting LLM-based Machine Translation to the Next Level
von: Feng, Zhaopeng, et al.
Veröffentlicht: (2024)
von: Feng, Zhaopeng, et al.
Veröffentlicht: (2024)
Obscured but Not Erased: Evaluating Nationality Bias in LLMs via Name-Based Bias Benchmarks
von: Pelosio, Giulio, et al.
Veröffentlicht: (2025)
von: Pelosio, Giulio, et al.
Veröffentlicht: (2025)
Trustworthy Social Bias Measurement
von: Bommasani, Rishi, et al.
Veröffentlicht: (2022)
von: Bommasani, Rishi, et al.
Veröffentlicht: (2022)
How Does Differential Privacy Affect Social Bias in LLMs? A Systematic Evaluation
von: Tenorio, Eduardo, et al.
Veröffentlicht: (2026)
von: Tenorio, Eduardo, et al.
Veröffentlicht: (2026)
Towards Proactive Personalization through Profile Customization for Individual Users in Dialogues
von: Zhang, Xiaotian, et al.
Veröffentlicht: (2025)
von: Zhang, Xiaotian, et al.
Veröffentlicht: (2025)
Are LLMs (Really) Ideological? An IRT-based Analysis and Alignment Tool for Perceived Socio-Economic Bias in LLMs
von: Wachter, Jasmin, et al.
Veröffentlicht: (2025)
von: Wachter, Jasmin, et al.
Veröffentlicht: (2025)
Implicit Bias in LLMs: A Survey
von: Lin, Xinru, et al.
Veröffentlicht: (2025)
von: Lin, Xinru, et al.
Veröffentlicht: (2025)
Capturing Bias Diversity in LLMs
von: Gosavi, Purva Prasad, et al.
Veröffentlicht: (2024)
von: Gosavi, Purva Prasad, et al.
Veröffentlicht: (2024)
Business as Rulesual: A Benchmark and Framework for Business Rule Flow Modeling with LLMs
von: Yang, Chen, et al.
Veröffentlicht: (2025)
von: Yang, Chen, et al.
Veröffentlicht: (2025)
Med-U1: Incentivizing Unified Medical Reasoning in LLMs via Large-scale Reinforcement Learning
von: Zhang, Xiaotian, et al.
Veröffentlicht: (2025)
von: Zhang, Xiaotian, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
BiasGuard: A Reasoning-enhanced Bias Detection Tool For Large Language Models
von: Fan, Zhiting, et al.
Veröffentlicht: (2025) -
FairMT-Bench: Benchmarking Fairness for Multi-turn Dialogue in Conversational LLMs
von: Fan, Zhiting, et al.
Veröffentlicht: (2024) -
Large Language Model Bias Mitigation from the Perspective of Knowledge Editing
von: Chen, Ruizhe, et al.
Veröffentlicht: (2024) -
Identifying and Mitigating Social Bias Knowledge in Language Models
von: Chen, Ruizhe, et al.
Veröffentlicht: (2024) -
FairSteer: Inference Time Debiasing for LLMs with Dynamic Activation Steering
von: Li, Yichen, et al.
Veröffentlicht: (2025)