Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks
Fuente:
arXiv
Guardado en:
| Autores principales: | Gugg, Regina, Niederländer, Selina, Stöckl, Andreas, Flechl, Martin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Exploring the Impact of Personality Traits on LLM Bias and Toxicity
por: Wang, Shuo, et al.
Publicado: (2025)
por: Wang, Shuo, et al.
Publicado: (2025)
Generative Exaggeration in LLM Social Agents: Consistency, Bias, and Toxicity
por: Nudo, Jacopo, et al.
Publicado: (2025)
por: Nudo, Jacopo, et al.
Publicado: (2025)
CyBiasBench: Benchmarking Bias in LLM Agents for Cyber-Attack Scenarios
por: Lim, Taein, et al.
Publicado: (2026)
por: Lim, Taein, et al.
Publicado: (2026)
Assessing Bias in Metric Models for LLM Open-Ended Generation Bias Benchmarks
por: Demchak, Nathaniel, et al.
Publicado: (2024)
por: Demchak, Nathaniel, et al.
Publicado: (2024)
Polyrating: A Cost-Effective and Bias-Aware Rating System for LLM Evaluation
por: Dekoninck, Jasper, et al.
Publicado: (2024)
por: Dekoninck, Jasper, et al.
Publicado: (2024)
Beyond Consensus: Mitigating the Agreeableness Bias in LLM Judge Evaluations
por: Jain, Suryaansh, et al.
Publicado: (2025)
por: Jain, Suryaansh, et al.
Publicado: (2025)
Investigating the Impact of LLM Personality on Cognitive Bias Manifestation in Automated Decision-Making Tasks
por: He, Jiangen, et al.
Publicado: (2025)
por: He, Jiangen, et al.
Publicado: (2025)
Fairness or Fluency? An Investigation into Language Bias of Pairwise LLM-as-a-Judge
por: Zhou, Xiaolin, et al.
Publicado: (2026)
por: Zhou, Xiaolin, et al.
Publicado: (2026)
Investigating Gender Bias in LLM-Generated Stories via Psychological Stereotypes
por: Masoudian, Shahed, et al.
Publicado: (2025)
por: Masoudian, Shahed, et al.
Publicado: (2025)
Towards Provably Unbiased LLM Judges via Bias-Bounded Evaluation
por: Feuer, Benjamin, et al.
Publicado: (2026)
por: Feuer, Benjamin, et al.
Publicado: (2026)
Persuasiveness and Bias in LLM: Investigating the Impact of Persuasiveness and Reinforcement of Bias in Language Models
por: Roy, Saumya
Publicado: (2025)
por: Roy, Saumya
Publicado: (2025)
Can LLMs Recognize Toxicity? A Structured Investigation Framework and Toxicity Metric
por: Koh, Hyukhun, et al.
Publicado: (2024)
por: Koh, Hyukhun, et al.
Publicado: (2024)
The influence of persona and conversational task on social interactions with a LLM-controlled embodied conversational agent
por: Kroczek, Leon O. H., et al.
Publicado: (2024)
por: Kroczek, Leon O. H., et al.
Publicado: (2024)
LLM Bias Evaluation: Gender, Racial, and Age Disparities in Occupational and Crime Scenarios
por: Mirza, Vishal, et al.
Publicado: (2024)
por: Mirza, Vishal, et al.
Publicado: (2024)
IndoBias: A Dual Track Culturally Grounded Benchmark for LLMs Bias Evaluation in Indonesian Languages
por: Hanif, Ikhlasul Akmal, et al.
Publicado: (2026)
por: Hanif, Ikhlasul Akmal, et al.
Publicado: (2026)
Social Bias in LLM-Generated Code: Benchmark and Mitigation
por: Rabbi, Fazle, et al.
Publicado: (2026)
por: Rabbi, Fazle, et al.
Publicado: (2026)
No LLM is Free From Bias: A Comprehensive Study of Bias Evaluation in Large Language Models
por: Kumar, Charaka Vinayak, et al.
Publicado: (2025)
por: Kumar, Charaka Vinayak, et al.
Publicado: (2025)
Does Reasoning Introduce Bias? A Study of Social Bias Evaluation and Mitigation in LLM Reasoning
por: Wu, Xuyang, et al.
Publicado: (2025)
por: Wu, Xuyang, et al.
Publicado: (2025)
When Wording Steers the Evaluation: Framing Bias in LLM judges
por: Hwang, Yerin, et al.
Publicado: (2026)
por: Hwang, Yerin, et al.
Publicado: (2026)
Evaluating and Mitigating LLM-as-a-judge Bias in Communication Systems
por: Gao, Jiaxin, et al.
Publicado: (2025)
por: Gao, Jiaxin, et al.
Publicado: (2025)
When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation
por: Xu, Wenda, et al.
Publicado: (2025)
por: Xu, Wenda, et al.
Publicado: (2025)
Enhancing LLM-based Hatred and Toxicity Detection with Meta-Toxic Knowledge Graph
por: Zhao, Yibo, et al.
Publicado: (2024)
por: Zhao, Yibo, et al.
Publicado: (2024)
Who Gets the Mic? Investigating Gender Bias in the Speaker Assignment of a Speech-LLM
por: Puhach, Dariia, et al.
Publicado: (2025)
por: Puhach, Dariia, et al.
Publicado: (2025)
BiasScope: Towards Automated Detection of Bias in LLM-as-a-Judge Evaluation
por: Lai, Peng, et al.
Publicado: (2026)
por: Lai, Peng, et al.
Publicado: (2026)
Identifying and Mitigating Systemic Measurement Bias in Production LLM Inference Benchmarks
por: Chandrasekar, Ashok, et al.
Publicado: (2026)
por: Chandrasekar, Ashok, et al.
Publicado: (2026)
Social Bias Benchmark for Generation: A Comparison of Generation and QA-Based Evaluations
por: Jin, Jiho, et al.
Publicado: (2025)
por: Jin, Jiho, et al.
Publicado: (2025)
Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines
por: Soumik, Sadman Kabir
Publicado: (2026)
por: Soumik, Sadman Kabir
Publicado: (2026)
The Evaluation Game: Beyond Static LLM Benchmarking
por: Wang, Paul, et al.
Publicado: (2026)
por: Wang, Paul, et al.
Publicado: (2026)
Evaluation and Benchmarking of LLM Agents: A Survey
por: Mohammadi, Mahmoud, et al.
Publicado: (2025)
por: Mohammadi, Mahmoud, et al.
Publicado: (2025)
Who Benchmarks the Benchmarks? A Case Study of LLM Evaluation in Icelandic
por: Ingimundarson, Finnur Ágúst, et al.
Publicado: (2026)
por: Ingimundarson, Finnur Ágúst, et al.
Publicado: (2026)
Mobile-Env: Building Qualified Evaluation Benchmarks for LLM-GUI Interaction
por: Zhang, Danyang, et al.
Publicado: (2023)
por: Zhang, Danyang, et al.
Publicado: (2023)
Social Evolution of Published Text and The Emergence of Artificial Intelligence Through Large Language Models and The Problem of Toxicity and Bias
por: Khan, Arifa, et al.
Publicado: (2024)
por: Khan, Arifa, et al.
Publicado: (2024)
LLM BiasScope: A Real-Time Bias Analysis Platform for Comparative LLM Evaluation
por: Ghosh, Himel, et al.
Publicado: (2026)
por: Ghosh, Himel, et al.
Publicado: (2026)
NC-Bench: An LLM Benchmark for Evaluating Conversational Competence
por: Moore, Robert J., et al.
Publicado: (2026)
por: Moore, Robert J., et al.
Publicado: (2026)
AI Benchmarks and Datasets for LLM Evaluation
por: Ivanov, Todor, et al.
Publicado: (2024)
por: Ivanov, Todor, et al.
Publicado: (2024)
RewardHackingAgents: Benchmarking Evaluation Integrity for LLM ML-Engineering Agents
por: Atinafu, Yonas, et al.
Publicado: (2026)
por: Atinafu, Yonas, et al.
Publicado: (2026)
BrowserArena: Evaluating LLM Agents on Real-World Web Navigation Tasks
por: Anupam, Sagnik, et al.
Publicado: (2025)
por: Anupam, Sagnik, et al.
Publicado: (2025)
Realistic Evaluation of Toxicity in Large Language Models
por: Luong, Tinh Son, et al.
Publicado: (2024)
por: Luong, Tinh Son, et al.
Publicado: (2024)
Benchmarking Practices in LLM-driven Offensive Security: Testbeds, Metrics, and Experiment Design
por: Happe, Andreas, et al.
Publicado: (2025)
por: Happe, Andreas, et al.
Publicado: (2025)
Reward Bias Substitution: Single-Axis Bias Mitigations Redirect Optimization Pressure
por: Lamparth, Max, et al.
Publicado: (2026)
por: Lamparth, Max, et al.
Publicado: (2026)
Ejemplares similares
-
Exploring the Impact of Personality Traits on LLM Bias and Toxicity
por: Wang, Shuo, et al.
Publicado: (2025) -
Generative Exaggeration in LLM Social Agents: Consistency, Bias, and Toxicity
por: Nudo, Jacopo, et al.
Publicado: (2025) -
CyBiasBench: Benchmarking Bias in LLM Agents for Cyber-Attack Scenarios
por: Lim, Taein, et al.
Publicado: (2026) -
Assessing Bias in Metric Models for LLM Open-Ended Generation Bias Benchmarks
por: Demchak, Nathaniel, et al.
Publicado: (2024) -
Polyrating: A Cost-Effective and Bias-Aware Rating System for LLM Evaluation
por: Dekoninck, Jasper, et al.
Publicado: (2024)