Saved in:
| Main Authors: | Nghiem, Huy, Gupta, Umang, Morstatter, Fred |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2402.03221 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HateCOT: An Explanation-Enhanced Dataset for Generalizable Offensive Speech Detection via Large Language Models
by: Nghiem, Huy, et al.
Published: (2024)
by: Nghiem, Huy, et al.
Published: (2024)
Harmful Speech Detection by Language Models Exhibits Gender-Queer Dialect Bias
by: Dorn, Rebecca, et al.
Published: (2024)
by: Dorn, Rebecca, et al.
Published: (2024)
The Butterfly Effect of Altering Prompts: How Small Changes and Jailbreaks Affect Large Language Model Performance
by: Salinas, Abel, et al.
Published: (2024)
by: Salinas, Abel, et al.
Published: (2024)
Risk and Response in Large Language Models: Evaluating Key Threat Categories
by: Harandizadeh, Bahareh, et al.
Published: (2024)
by: Harandizadeh, Bahareh, et al.
Published: (2024)
Knowledge Graph Analysis of Legal Understanding and Violations in LLMs
by: Jha, Abha, et al.
Published: (2025)
by: Jha, Abha, et al.
Published: (2025)
Do Methods to Jailbreak and Defend LLMs Generalize Across Languages?
by: Atil, Berk, et al.
Published: (2025)
by: Atil, Berk, et al.
Published: (2025)
Contextualizing Argument Quality Assessment with Relevant Knowledge
by: Deshpande, Darshan, et al.
Published: (2023)
by: Deshpande, Darshan, et al.
Published: (2023)
Leveraging Sentiment for Offensive Text Classification
by: Islam, Khondoker Ittehadul
Published: (2024)
by: Islam, Khondoker Ittehadul
Published: (2024)
Estimating Causal Effects of Text Interventions Leveraging LLMs
by: Guo, Siyi, et al.
Published: (2024)
by: Guo, Siyi, et al.
Published: (2024)
SMARTER: A Data-efficient Framework to Improve Toxicity Detection with Explanation via Self-augmenting Large Language Models
by: Nghiem, Huy, et al.
Published: (2025)
by: Nghiem, Huy, et al.
Published: (2025)
Vicarious Offense and Noise Audit of Offensive Speech Classifiers: Unifying Human and Machine Disagreement on What is Offensive
by: Weerasooriya, Tharindu Cyril, et al.
Published: (2023)
by: Weerasooriya, Tharindu Cyril, et al.
Published: (2023)
Multi-task Learning with Active Learning for Arabic Offensive Speech Detection
by: Alansari, Aisha, et al.
Published: (2025)
by: Alansari, Aisha, et al.
Published: (2025)
"You Gotta be a Doctor, Lin": An Investigation of Name-Based Bias of Large Language Models in Employment Recommendations
by: Nghiem, Huy, et al.
Published: (2024)
by: Nghiem, Huy, et al.
Published: (2024)
The Unequal Opportunities of Large Language Models: Revealing Demographic Bias through Job Recommendations
by: Salinas, Abel, et al.
Published: (2023)
by: Salinas, Abel, et al.
Published: (2023)
OffensiveLang: A Community Based Implicit Offensive Language Dataset
by: Das, Amit, et al.
Published: (2024)
by: Das, Amit, et al.
Published: (2024)
Balancing Safety and Helpfulness in Healthcare AI Assistants through Iterative Preference Alignment
by: Nghiem, Huy, et al.
Published: (2025)
by: Nghiem, Huy, et al.
Published: (2025)
CFMatch: Aligning Automated Answer Equivalence Evaluation with Expert Judgments For Open-Domain Question Answering
by: Li, Zongxia, et al.
Published: (2024)
by: Li, Zongxia, et al.
Published: (2024)
The Curious Case of Nonverbal Abstract Reasoning with Multi-Modal Large Language Models
by: Ahrabian, Kian, et al.
Published: (2024)
by: Ahrabian, Kian, et al.
Published: (2024)
Capturing Perspectives of Crowdsourced Annotators in Subjective Learning Tasks
by: Mokhberian, Negar, et al.
Published: (2023)
by: Mokhberian, Negar, et al.
Published: (2023)
Refining Input Guardrails: Enhancing LLM-as-a-Judge Efficiency Through Chain-of-Thought Fine-Tuning and Alignment
by: Rad, Melissa Kazemi, et al.
Published: (2025)
by: Rad, Melissa Kazemi, et al.
Published: (2025)
Exploring Boundaries and Intensities in Offensive and Hate Speech: Unveiling the Complex Spectrum of Social Media Discourse
by: Ayele, Abinew Ali, et al.
Published: (2024)
by: Ayele, Abinew Ali, et al.
Published: (2024)
'Rich Dad, Poor Lad': How do Large Language Models Contextualize Socioeconomic Factors in College Admission ?
by: Nghiem, Huy, et al.
Published: (2025)
by: Nghiem, Huy, et al.
Published: (2025)
Can You Make It Sound Like You? Post-Editing LLM-Generated Text for Personal Style
by: Baumler, Connor, et al.
Published: (2026)
by: Baumler, Connor, et al.
Published: (2026)
Bias in the Tails: How Name-conditioned Evaluative Framing in Resume Summaries Destabilizes LLM-based Hiring
by: Nghiem, Huy, et al.
Published: (2026)
by: Nghiem, Huy, et al.
Published: (2026)
A Modular Taxonomy for Hate Speech Definitions and Its Impact on Zero-Shot LLM Classification Performance
by: Melis, Matteo, et al.
Published: (2025)
by: Melis, Matteo, et al.
Published: (2025)
Enhanced Data Race Prediction Through Modular Reasoning
by: Ang, Zhendong, et al.
Published: (2025)
by: Ang, Zhendong, et al.
Published: (2025)
PEDANTS: Cheap but Effective and Interpretable Answer Equivalence
by: Li, Zongxia, et al.
Published: (2024)
by: Li, Zongxia, et al.
Published: (2024)
Don't Blame the Data, Blame the Model: Understanding Noise and Bias When Learning from Subjective Annotations
by: Anand, Abhishek, et al.
Published: (2024)
by: Anand, Abhishek, et al.
Published: (2024)
Secret Keepers: The Impact of LLMs on Linguistic Markers of Personal Traits
by: Sourati, Zhivar, et al.
Published: (2024)
by: Sourati, Zhivar, et al.
Published: (2024)
Reinforcing Stereotypes of Anger: Emotion AI on African American Vernacular English
by: Dorn, Rebecca, et al.
Published: (2025)
by: Dorn, Rebecca, et al.
Published: (2025)
Efficient Decrease-and-Conquer Linearizability Monitoring
by: Han, Lee Zheng, et al.
Published: (2024)
by: Han, Lee Zheng, et al.
Published: (2024)
Towards Generalized Offensive Language Identification
by: Dmonte, Alphaeus, et al.
Published: (2024)
by: Dmonte, Alphaeus, et al.
Published: (2024)
Parametrizing Reads-From Equivalence for Predictive Monitoring
by: Farzan, Azadeh, et al.
Published: (2026)
by: Farzan, Azadeh, et al.
Published: (2026)
Offset Unlearning for Large Language Models
by: Huang, James Y., et al.
Published: (2024)
by: Huang, James Y., et al.
Published: (2024)
Enhancing Robustness of AI Offensive Code Generators via Data Augmentation
by: Improta, Cristina, et al.
Published: (2023)
by: Improta, Cristina, et al.
Published: (2023)
Systematic Offensive Stereotyping (SOS) Bias in Language Models
by: Elsafoury, Fatma
Published: (2023)
by: Elsafoury, Fatma
Published: (2023)
Detection and Analysis of Offensive Online Content in Hausa Language
by: Adam, Fatima Muhammad, et al.
Published: (2023)
by: Adam, Fatima Muhammad, et al.
Published: (2023)
A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges
by: Li, Zongxia, et al.
Published: (2025)
by: Li, Zongxia, et al.
Published: (2025)
Enhancing Romanian Offensive Language Detection through Knowledge Distillation, Multi-Task Learning, and Data Augmentation
by: Matei, Vlad-Cristian, et al.
Published: (2024)
by: Matei, Vlad-Cristian, et al.
Published: (2024)
Robustness and Confounders in the Demographic Alignment of LLMs with Human Perceptions of Offensiveness
by: Alipour, Shayan, et al.
Published: (2024)
by: Alipour, Shayan, et al.
Published: (2024)
Similar Items
-
HateCOT: An Explanation-Enhanced Dataset for Generalizable Offensive Speech Detection via Large Language Models
by: Nghiem, Huy, et al.
Published: (2024) -
Harmful Speech Detection by Language Models Exhibits Gender-Queer Dialect Bias
by: Dorn, Rebecca, et al.
Published: (2024) -
The Butterfly Effect of Altering Prompts: How Small Changes and Jailbreaks Affect Large Language Model Performance
by: Salinas, Abel, et al.
Published: (2024) -
Risk and Response in Large Language Models: Evaluating Key Threat Categories
by: Harandizadeh, Bahareh, et al.
Published: (2024) -
Knowledge Graph Analysis of Legal Understanding and Violations in LLMs
by: Jha, Abha, et al.
Published: (2025)