Down the Toxicity Rabbit Hole: A Novel Framework to Bias Audit Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Dutta, Arka, Khorramrouz, Adel, Dutta, Sujan, KhudaBukhsh, Ashiqur R. |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Gender Representation and Bias in Indian Civil Service Mock Interviews
by: Banerjee, Somonnoy, et al.
Published: (2024)
by: Banerjee, Somonnoy, et al.
Published: (2024)
What About the Scene with the Hitler Reference? HAUNT: A Framework to Probe LLMs' Self-consistency Via Adversarial Nudge
by: Dutta, Arka, et al.
Published: (2025)
by: Dutta, Arka, et al.
Published: (2025)
Navigating the Rabbit Hole: Emergent Biases in LLM-Generated Attack Narratives Targeting Mental Health Groups
by: Magu, Rijul, et al.
Published: (2025)
by: Magu, Rijul, et al.
Published: (2025)
Vicarious Offense and Noise Audit of Offensive Speech Classifiers: Unifying Human and Machine Disagreement on What is Offensive
by: Weerasooriya, Tharindu Cyril, et al.
Published: (2023)
by: Weerasooriya, Tharindu Cyril, et al.
Published: (2023)
Characterizing Selective Refusal Bias in Large Language Models
by: Khorramrouz, Adel, et al.
Published: (2025)
by: Khorramrouz, Adel, et al.
Published: (2025)
Community Needs and Assets: A Computational Analysis of Community Conversations
by: Chowdhury, Md Towhidul Absar, et al.
Published: (2024)
by: Chowdhury, Md Towhidul Absar, et al.
Published: (2024)
Hope vs. Hate: Understanding User Interactions with LGBTQ+ News Content in Mainstream US News Media through the Lens of Hope Speech
by: Pofcher, Jonathan, et al.
Published: (2025)
by: Pofcher, Jonathan, et al.
Published: (2025)
When Neutral Summaries are not that Neutral: Quantifying Political Neutrality in LLM-Generated News Summaries
by: Vijay, Supriti, et al.
Published: (2024)
by: Vijay, Supriti, et al.
Published: (2024)
Infrastructure Ombudsman: Mining Future Failure Concerns from Structural Disaster Response
by: Chowdhury, Md Towhidul Absar, et al.
Published: (2024)
by: Chowdhury, Md Towhidul Absar, et al.
Published: (2024)
ARTICLE: Annotator Reliability Through In-Context Learning
by: Dutta, Sujan, et al.
Published: (2024)
by: Dutta, Sujan, et al.
Published: (2024)
Investigating Vaccine Buyer's Remorse: Post-Vaccination Decision Regret in COVID-19 Social Media Using Politically Diverse Human Annotation
by: Stanley, Miles, et al.
Published: (2026)
by: Stanley, Miles, et al.
Published: (2026)
Rater Cohesion and Quality from a Vicarious Perspective
by: Pandita, Deepak, et al.
Published: (2024)
by: Pandita, Deepak, et al.
Published: (2024)
Mapping Violence: Developing an Extensive Framework to Build a Bangla Sectarian Expression Dataset from Social Media Interactions
by: Tasnim, Nazia, et al.
Published: (2024)
by: Tasnim, Nazia, et al.
Published: (2024)
Datasets for Depression Modeling in Social Media: An Overview
by: Bucur, Ana-Maria, et al.
Published: (2025)
by: Bucur, Ana-Maria, et al.
Published: (2025)
On the State of NLP Approaches to Modeling Depression in Social Media: A Post-COVID-19 Outlook
by: Bucur, Ana-Maria, et al.
Published: (2024)
by: Bucur, Ana-Maria, et al.
Published: (2024)
Online Anti-sexist Speech: Identifying Resistance to Gender Bias in Political Discourse
by: Dutta, Aditi, et al.
Published: (2025)
by: Dutta, Aditi, et al.
Published: (2025)
Political Alignment in Large Language Models: A Multidimensional Audit of Psychometric Identity and Behavioral Bias
by: Sakhawat, Adib, et al.
Published: (2026)
by: Sakhawat, Adib, et al.
Published: (2026)
What's in a Name? Auditing Large Language Models for Race and Gender Bias
by: Salinas, Alejandro, et al.
Published: (2024)
by: Salinas, Alejandro, et al.
Published: (2024)
AuditWen:An Open-Source Large Language Model for Audit
by: Huang, Jiajia, et al.
Published: (2024)
by: Huang, Jiajia, et al.
Published: (2024)
Toward Socially Aware Vision-Language Models: Evaluating Cultural Competence Through Multimodal Story Generation
by: Mukherjee, Arka, et al.
Published: (2025)
by: Mukherjee, Arka, et al.
Published: (2025)
Walking in Others' Shoes: How Perspective-Taking Guides Large Language Models in Reducing Toxicity and Bias
by: Xu, Rongwu, et al.
Published: (2024)
by: Xu, Rongwu, et al.
Published: (2024)
Gender Bias in Emotion Recognition by Large Language Models
by: Herbert, Maureen, et al.
Published: (2025)
by: Herbert, Maureen, et al.
Published: (2025)
Bias and Volatility: A Statistical Framework for Evaluating Large Language Model's Stereotypes and the Associated Generation Inconsistency
by: Liu, Yiran, et al.
Published: (2024)
by: Liu, Yiran, et al.
Published: (2024)
Measuring Implicit Bias in Explicitly Unbiased Large Language Models
by: Bai, Xuechunzi, et al.
Published: (2024)
by: Bai, Xuechunzi, et al.
Published: (2024)
Including frameworks of public health ethics in computational modelling of infectious disease interventions
by: Zarebski, Alexander E., et al.
Published: (2025)
by: Zarebski, Alexander E., et al.
Published: (2025)
When AI Speaks, Whose Values Does It Express? A Cross-Cultural Audit of Individualism-Collectivism Bias in Large Language Models
by: Venkata, Pruthvinath Jeripity
Published: (2026)
by: Venkata, Pruthvinath Jeripity
Published: (2026)
PRISM: A Methodology for Auditing Biases in Large Language Models
by: Azzopardi, Leif, et al.
Published: (2024)
by: Azzopardi, Leif, et al.
Published: (2024)
Leveraging Large Language Models to Measure Gender Representation Bias in Gendered Language Corpora
by: Derner, Erik, et al.
Published: (2024)
by: Derner, Erik, et al.
Published: (2024)
Cross-Language Bias Examination in Large Language Models
by: Liang, Yuxuan, et al.
Published: (2025)
by: Liang, Yuxuan, et al.
Published: (2025)
JobFair: A Framework for Benchmarking Gender Hiring Bias in Large Language Models
by: Wang, Ze, et al.
Published: (2024)
by: Wang, Ze, et al.
Published: (2024)
Towards Equitable AI: Detecting Bias in Using Large Language Models for Marketing
by: Yilmaz, Berk, et al.
Published: (2025)
by: Yilmaz, Berk, et al.
Published: (2025)
Characterizing Bias: Benchmarking Large Language Models in Simplified versus Traditional Chinese
by: Lyu, Hanjia, et al.
Published: (2025)
by: Lyu, Hanjia, et al.
Published: (2025)
Divided by discipline? A systematic literature review on the quantification of online sexism and misogyny using a semi-automated approach
by: Dutta, Aditi, et al.
Published: (2024)
by: Dutta, Aditi, et al.
Published: (2024)
Inference-Time Reasoning Selectively Reduces Implicit Social Bias in Large Language Models
by: Apsel, Molly, et al.
Published: (2026)
by: Apsel, Molly, et al.
Published: (2026)
Translate With Care: Addressing Gender Bias, Neutrality, and Reasoning in Large Language Model Translations
by: Zahraei, Pardis Sadat, et al.
Published: (2025)
by: Zahraei, Pardis Sadat, et al.
Published: (2025)
WinoQueer: A Community-in-the-Loop Benchmark for Anti-LGBTQ+ Bias in Large Language Models
by: Felkner, Virginia K., et al.
Published: (2023)
by: Felkner, Virginia K., et al.
Published: (2023)
Divine LLaMAs: Bias, Stereotypes, Stigmatization, and Emotion Representation of Religion in Large Language Models
by: Plaza-del-Arco, Flor Miriam, et al.
Published: (2024)
by: Plaza-del-Arco, Flor Miriam, et al.
Published: (2024)
DECASTE: Unveiling Caste Stereotypes in Large Language Models through Multi-Dimensional Bias Analysis
by: Vijayaraghavan, Prashanth, et al.
Published: (2025)
by: Vijayaraghavan, Prashanth, et al.
Published: (2025)
Examining Gender and Racial Bias in Large Vision-Language Models Using a Novel Dataset of Parallel Images
by: Fraser, Kathleen C., et al.
Published: (2024)
by: Fraser, Kathleen C., et al.
Published: (2024)
Information Suppression in Large Language Models: Auditing, Quantifying, and Characterizing Censorship in DeepSeek
by: Qiu, Peiran, et al.
Published: (2025)
by: Qiu, Peiran, et al.
Published: (2025)
Similar Items
-
Gender Representation and Bias in Indian Civil Service Mock Interviews
by: Banerjee, Somonnoy, et al.
Published: (2024) -
What About the Scene with the Hitler Reference? HAUNT: A Framework to Probe LLMs' Self-consistency Via Adversarial Nudge
by: Dutta, Arka, et al.
Published: (2025) -
Navigating the Rabbit Hole: Emergent Biases in LLM-Generated Attack Narratives Targeting Mental Health Groups
by: Magu, Rijul, et al.
Published: (2025) -
Vicarious Offense and Noise Audit of Offensive Speech Classifiers: Unifying Human and Machine Disagreement on What is Offensive
by: Weerasooriya, Tharindu Cyril, et al.
Published: (2023) -
Characterizing Selective Refusal Bias in Large Language Models
by: Khorramrouz, Adel, et al.
Published: (2025)