From Perceived Effectiveness to Measured Impact: Identity-Aware Evaluation of Automated Counter-Stereotypes
Fuente:
arXiv
Saved in:
| Main Authors: | Kiritchenko, Svetlana, Kerkhof, Anna, Nejadgholi, Isar, Fraser, Kathleen C. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Challenging Negative Gender Stereotypes: A Study on the Effectiveness of Automated Counter-Stereotypes
by: Nejadgholi, Isar, et al.
Published: (2024)
by: Nejadgholi, Isar, et al.
Published: (2024)
Tackling Social Bias against the Poor: A Dataset and Taxonomy on Aporophobia
by: Curto, Georgina, et al.
Published: (2025)
by: Curto, Georgina, et al.
Published: (2025)
Fine-Tuning Lowers Safety and Disrupts Evaluation Consistency
by: Fraser, Kathleen C., et al.
Published: (2025)
by: Fraser, Kathleen C., et al.
Published: (2025)
The crime of being poor
by: Curto, Georgina, et al.
Published: (2023)
by: Curto, Georgina, et al.
Published: (2023)
Examining Gender and Racial Bias in Large Vision-Language Models Using a Novel Dataset of Parallel Images
by: Fraser, Kathleen C., et al.
Published: (2024)
by: Fraser, Kathleen C., et al.
Published: (2024)
Adaptable Moral Stances of Large Language Models on Sexist Content: Implications for Society and Gender Discourse
by: Guo, Rongchen, et al.
Published: (2024)
by: Guo, Rongchen, et al.
Published: (2024)
Detecting AI-Generated Text: Factors Influencing Detectability with Current Methods
by: Fraser, Kathleen C., et al.
Published: (2024)
by: Fraser, Kathleen C., et al.
Published: (2024)
Human-Centered AI Applications for Canada's Immigration Settlement Sector
by: Nejadgholi, Isar, et al.
Published: (2024)
by: Nejadgholi, Isar, et al.
Published: (2024)
Social and Ethical Risks Posed by General-Purpose LLMs for Settling Newcomers in Canada
by: Nejadgholi, Isar, et al.
Published: (2024)
by: Nejadgholi, Isar, et al.
Published: (2024)
When Detection Fails: The Power of Fine-Tuned Models to Generate Human-Like Social Media Text
by: Dawkins, Hillary, et al.
Published: (2025)
by: Dawkins, Hillary, et al.
Published: (2025)
Uncovering Bias in Large Vision-Language Models at Scale with Counterfactuals
by: Howard, Phillip, et al.
Published: (2024)
by: Howard, Phillip, et al.
Published: (2024)
Detecting Gender Stereotypes in Scratch Programming Tutorials
by: Graßl, Isabella, et al.
Published: (2025)
by: Graßl, Isabella, et al.
Published: (2025)
Uncovering Bias in Large Vision-Language Models with Counterfactuals
by: Howard, Phillip, et al.
Published: (2024)
by: Howard, Phillip, et al.
Published: (2024)
Socially Aware Synthetic Data Generation for Suicidal Ideation Detection Using Large Language Models
by: Ghanadian, Hamideh, et al.
Published: (2024)
by: Ghanadian, Hamideh, et al.
Published: (2024)
Defining Cultural Capabilities for AI Evaluation: A Taxonomy Grounded in Intercultural Communication Theory
by: Nejadgholi, Isar, et al.
Published: (2026)
by: Nejadgholi, Isar, et al.
Published: (2026)
SESGO: Spanish Evaluation of Stereotypical Generative Outputs
by: Robles, Melissa, et al.
Published: (2025)
by: Robles, Melissa, et al.
Published: (2025)
Gender-Neutral Machine Translation Strategies in Practice
by: Dawkins, Hillary, et al.
Published: (2025)
by: Dawkins, Hillary, et al.
Published: (2025)
WMT24 Test Suite: Gender Resolution in Speaker-Listener Dialogue Roles
by: Dawkins, Hillary, et al.
Published: (2024)
by: Dawkins, Hillary, et al.
Published: (2024)
A Taxonomy for Design and Evaluation of Prompt-Based Natural Language Explanations
by: Nejadgholi, Isar, et al.
Published: (2025)
by: Nejadgholi, Isar, et al.
Published: (2025)
A Comprehensive Framework to Operationalize Social Stereotypes for Responsible AI Evaluations
by: Davani, Aida, et al.
Published: (2025)
by: Davani, Aida, et al.
Published: (2025)
Multilingual Text-to-Image Generation Magnifies Gender Stereotypes and Prompt Engineering May Not Help You
by: Friedrich, Felix, et al.
Published: (2024)
by: Friedrich, Felix, et al.
Published: (2024)
Evaluating Chinese Large Language Models: The Influence of Persona Assignment on Stereotypes and Safeguards
by: Liu, Geng, et al.
Published: (2025)
by: Liu, Geng, et al.
Published: (2025)
Adaptive Data Collection for Latin-American Community-sourced Evaluation of Stereotypes (LACES)
by: Ivetta, Guido, et al.
Published: (2025)
by: Ivetta, Guido, et al.
Published: (2025)
Automated Transparency: A Legal and Empirical Analysis of the Digital Services Act Transparency Database
by: Kaushal, Rishabh, et al.
Published: (2024)
by: Kaushal, Rishabh, et al.
Published: (2024)
Surfacing Subtle Stereotypes: A Multilingual, Debate-Oriented Evaluation of Modern LLMs
by: Saeed, Muhammed, et al.
Published: (2025)
by: Saeed, Muhammed, et al.
Published: (2025)
LLM Judges Inconsistently Disagree Across Safety Criteria and Harm Categories
by: Vishnubhotla, Krishnapriya, et al.
Published: (2026)
by: Vishnubhotla, Krishnapriya, et al.
Published: (2026)
Local Contrastive Editing of Gender Stereotypes
by: Lutz, Marlene, et al.
Published: (2024)
by: Lutz, Marlene, et al.
Published: (2024)
Bias and Volatility: A Statistical Framework for Evaluating Large Language Model's Stereotypes and the Associated Generation Inconsistency
by: Liu, Yiran, et al.
Published: (2024)
by: Liu, Yiran, et al.
Published: (2024)
Examining Racial Stereotypes in YouTube Autocomplete Suggestions
by: Ha, Eunbin, et al.
Published: (2024)
by: Ha, Eunbin, et al.
Published: (2024)
Projective Methods for Mitigating Gender Bias in Pre-trained Language Models
by: Dawkins, Hillary, et al.
Published: (2024)
by: Dawkins, Hillary, et al.
Published: (2024)
A Survey on Stereotype Detection in Natural Language Processing
by: Cignarella, Alessandra Teresa, et al.
Published: (2025)
by: Cignarella, Alessandra Teresa, et al.
Published: (2025)
Interpretations, Representations, and Stereotypes of Caste within Text-to-Image Generators
by: Ghosh, Sourojit
Published: (2024)
by: Ghosh, Sourojit
Published: (2024)
Countering Privacy Nihilism
by: Engelmann, Severin, et al.
Published: (2025)
by: Engelmann, Severin, et al.
Published: (2025)
Measuring Machine Learning Harms from Stereotypes Requires Understanding Who Is Harmed by Which Errors in What Ways
by: Wang, Angelina, et al.
Published: (2024)
by: Wang, Angelina, et al.
Published: (2024)
$\texttt{ModSCAN}$: Measuring Stereotypical Bias in Large Vision-Language Models from Vision and Language Modalities
by: Jiang, Yukun, et al.
Published: (2024)
by: Jiang, Yukun, et al.
Published: (2024)
Perceiving and Countering Hate: The Role of Identity in Online Responses
by: Ping, Kaike, et al.
Published: (2024)
by: Ping, Kaike, et al.
Published: (2024)
Measuring User Perceived Security of Mobile Banking Applications
by: Apaua, Richard, et al.
Published: (2022)
by: Apaua, Richard, et al.
Published: (2022)
Quantifying Gender Stereotypes in Japan between 1900 and 1999 with Word Embeddings
by: Sakai, Shintaro, et al.
Published: (2025)
by: Sakai, Shintaro, et al.
Published: (2025)
An LLM-based Chain-of-Response Counter-Scam System
by: Kim, Heedou, et al.
Published: (2026)
by: Kim, Heedou, et al.
Published: (2026)
Analyzing the Safety of Japanese Large Language Models in Stereotype-Triggering Prompts
by: Nakanishi, Akito, et al.
Published: (2025)
by: Nakanishi, Akito, et al.
Published: (2025)
Similar Items
-
Challenging Negative Gender Stereotypes: A Study on the Effectiveness of Automated Counter-Stereotypes
by: Nejadgholi, Isar, et al.
Published: (2024) -
Tackling Social Bias against the Poor: A Dataset and Taxonomy on Aporophobia
by: Curto, Georgina, et al.
Published: (2025) -
Fine-Tuning Lowers Safety and Disrupts Evaluation Consistency
by: Fraser, Kathleen C., et al.
Published: (2025) -
The crime of being poor
by: Curto, Georgina, et al.
Published: (2023) -
Examining Gender and Racial Bias in Large Vision-Language Models Using a Novel Dataset of Parallel Images
by: Fraser, Kathleen C., et al.
Published: (2024)