Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Gugg, Regina, Niederländer, Selina, Stöckl, Andreas, Flechl, Martin |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Exploring the Impact of Personality Traits on LLM Bias and Toxicity
par: Wang, Shuo, et autres
Publié: (2025)
par: Wang, Shuo, et autres
Publié: (2025)
Generative Exaggeration in LLM Social Agents: Consistency, Bias, and Toxicity
par: Nudo, Jacopo, et autres
Publié: (2025)
par: Nudo, Jacopo, et autres
Publié: (2025)
CyBiasBench: Benchmarking Bias in LLM Agents for Cyber-Attack Scenarios
par: Lim, Taein, et autres
Publié: (2026)
par: Lim, Taein, et autres
Publié: (2026)
Assessing Bias in Metric Models for LLM Open-Ended Generation Bias Benchmarks
par: Demchak, Nathaniel, et autres
Publié: (2024)
par: Demchak, Nathaniel, et autres
Publié: (2024)
Polyrating: A Cost-Effective and Bias-Aware Rating System for LLM Evaluation
par: Dekoninck, Jasper, et autres
Publié: (2024)
par: Dekoninck, Jasper, et autres
Publié: (2024)
Beyond Consensus: Mitigating the Agreeableness Bias in LLM Judge Evaluations
par: Jain, Suryaansh, et autres
Publié: (2025)
par: Jain, Suryaansh, et autres
Publié: (2025)
Investigating the Impact of LLM Personality on Cognitive Bias Manifestation in Automated Decision-Making Tasks
par: He, Jiangen, et autres
Publié: (2025)
par: He, Jiangen, et autres
Publié: (2025)
Fairness or Fluency? An Investigation into Language Bias of Pairwise LLM-as-a-Judge
par: Zhou, Xiaolin, et autres
Publié: (2026)
par: Zhou, Xiaolin, et autres
Publié: (2026)
Investigating Gender Bias in LLM-Generated Stories via Psychological Stereotypes
par: Masoudian, Shahed, et autres
Publié: (2025)
par: Masoudian, Shahed, et autres
Publié: (2025)
Towards Provably Unbiased LLM Judges via Bias-Bounded Evaluation
par: Feuer, Benjamin, et autres
Publié: (2026)
par: Feuer, Benjamin, et autres
Publié: (2026)
Persuasiveness and Bias in LLM: Investigating the Impact of Persuasiveness and Reinforcement of Bias in Language Models
par: Roy, Saumya
Publié: (2025)
par: Roy, Saumya
Publié: (2025)
Can LLMs Recognize Toxicity? A Structured Investigation Framework and Toxicity Metric
par: Koh, Hyukhun, et autres
Publié: (2024)
par: Koh, Hyukhun, et autres
Publié: (2024)
The influence of persona and conversational task on social interactions with a LLM-controlled embodied conversational agent
par: Kroczek, Leon O. H., et autres
Publié: (2024)
par: Kroczek, Leon O. H., et autres
Publié: (2024)
LLM Bias Evaluation: Gender, Racial, and Age Disparities in Occupational and Crime Scenarios
par: Mirza, Vishal, et autres
Publié: (2024)
par: Mirza, Vishal, et autres
Publié: (2024)
IndoBias: A Dual Track Culturally Grounded Benchmark for LLMs Bias Evaluation in Indonesian Languages
par: Hanif, Ikhlasul Akmal, et autres
Publié: (2026)
par: Hanif, Ikhlasul Akmal, et autres
Publié: (2026)
Social Bias in LLM-Generated Code: Benchmark and Mitigation
par: Rabbi, Fazle, et autres
Publié: (2026)
par: Rabbi, Fazle, et autres
Publié: (2026)
No LLM is Free From Bias: A Comprehensive Study of Bias Evaluation in Large Language Models
par: Kumar, Charaka Vinayak, et autres
Publié: (2025)
par: Kumar, Charaka Vinayak, et autres
Publié: (2025)
Does Reasoning Introduce Bias? A Study of Social Bias Evaluation and Mitigation in LLM Reasoning
par: Wu, Xuyang, et autres
Publié: (2025)
par: Wu, Xuyang, et autres
Publié: (2025)
When Wording Steers the Evaluation: Framing Bias in LLM judges
par: Hwang, Yerin, et autres
Publié: (2026)
par: Hwang, Yerin, et autres
Publié: (2026)
Evaluating and Mitigating LLM-as-a-judge Bias in Communication Systems
par: Gao, Jiaxin, et autres
Publié: (2025)
par: Gao, Jiaxin, et autres
Publié: (2025)
When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation
par: Xu, Wenda, et autres
Publié: (2025)
par: Xu, Wenda, et autres
Publié: (2025)
Enhancing LLM-based Hatred and Toxicity Detection with Meta-Toxic Knowledge Graph
par: Zhao, Yibo, et autres
Publié: (2024)
par: Zhao, Yibo, et autres
Publié: (2024)
Who Gets the Mic? Investigating Gender Bias in the Speaker Assignment of a Speech-LLM
par: Puhach, Dariia, et autres
Publié: (2025)
par: Puhach, Dariia, et autres
Publié: (2025)
BiasScope: Towards Automated Detection of Bias in LLM-as-a-Judge Evaluation
par: Lai, Peng, et autres
Publié: (2026)
par: Lai, Peng, et autres
Publié: (2026)
Identifying and Mitigating Systemic Measurement Bias in Production LLM Inference Benchmarks
par: Chandrasekar, Ashok, et autres
Publié: (2026)
par: Chandrasekar, Ashok, et autres
Publié: (2026)
Social Bias Benchmark for Generation: A Comparison of Generation and QA-Based Evaluations
par: Jin, Jiho, et autres
Publié: (2025)
par: Jin, Jiho, et autres
Publié: (2025)
Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines
par: Soumik, Sadman Kabir
Publié: (2026)
par: Soumik, Sadman Kabir
Publié: (2026)
The Evaluation Game: Beyond Static LLM Benchmarking
par: Wang, Paul, et autres
Publié: (2026)
par: Wang, Paul, et autres
Publié: (2026)
Evaluation and Benchmarking of LLM Agents: A Survey
par: Mohammadi, Mahmoud, et autres
Publié: (2025)
par: Mohammadi, Mahmoud, et autres
Publié: (2025)
Who Benchmarks the Benchmarks? A Case Study of LLM Evaluation in Icelandic
par: Ingimundarson, Finnur Ágúst, et autres
Publié: (2026)
par: Ingimundarson, Finnur Ágúst, et autres
Publié: (2026)
Mobile-Env: Building Qualified Evaluation Benchmarks for LLM-GUI Interaction
par: Zhang, Danyang, et autres
Publié: (2023)
par: Zhang, Danyang, et autres
Publié: (2023)
Social Evolution of Published Text and The Emergence of Artificial Intelligence Through Large Language Models and The Problem of Toxicity and Bias
par: Khan, Arifa, et autres
Publié: (2024)
par: Khan, Arifa, et autres
Publié: (2024)
LLM BiasScope: A Real-Time Bias Analysis Platform for Comparative LLM Evaluation
par: Ghosh, Himel, et autres
Publié: (2026)
par: Ghosh, Himel, et autres
Publié: (2026)
NC-Bench: An LLM Benchmark for Evaluating Conversational Competence
par: Moore, Robert J., et autres
Publié: (2026)
par: Moore, Robert J., et autres
Publié: (2026)
AI Benchmarks and Datasets for LLM Evaluation
par: Ivanov, Todor, et autres
Publié: (2024)
par: Ivanov, Todor, et autres
Publié: (2024)
RewardHackingAgents: Benchmarking Evaluation Integrity for LLM ML-Engineering Agents
par: Atinafu, Yonas, et autres
Publié: (2026)
par: Atinafu, Yonas, et autres
Publié: (2026)
BrowserArena: Evaluating LLM Agents on Real-World Web Navigation Tasks
par: Anupam, Sagnik, et autres
Publié: (2025)
par: Anupam, Sagnik, et autres
Publié: (2025)
Realistic Evaluation of Toxicity in Large Language Models
par: Luong, Tinh Son, et autres
Publié: (2024)
par: Luong, Tinh Son, et autres
Publié: (2024)
Benchmarking Practices in LLM-driven Offensive Security: Testbeds, Metrics, and Experiment Design
par: Happe, Andreas, et autres
Publié: (2025)
par: Happe, Andreas, et autres
Publié: (2025)
Reward Bias Substitution: Single-Axis Bias Mitigations Redirect Optimization Pressure
par: Lamparth, Max, et autres
Publié: (2026)
par: Lamparth, Max, et autres
Publié: (2026)
Documents similaires
-
Exploring the Impact of Personality Traits on LLM Bias and Toxicity
par: Wang, Shuo, et autres
Publié: (2025) -
Generative Exaggeration in LLM Social Agents: Consistency, Bias, and Toxicity
par: Nudo, Jacopo, et autres
Publié: (2025) -
CyBiasBench: Benchmarking Bias in LLM Agents for Cyber-Attack Scenarios
par: Lim, Taein, et autres
Publié: (2026) -
Assessing Bias in Metric Models for LLM Open-Ended Generation Bias Benchmarks
par: Demchak, Nathaniel, et autres
Publié: (2024) -
Polyrating: A Cost-Effective and Bias-Aware Rating System for LLM Evaluation
par: Dekoninck, Jasper, et autres
Publié: (2024)