Investigating Annotator Bias in Large Language Models for Hate Speech Detection
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Das, Amit, Zhang, Zheng, Hasan, Najib, Sarkar, Souvika, Jamshidi, Fatemeh, Bhattacharya, Tathagata, Rahgouy, Mostafa, Raychawdhary, Nilanjana, Feng, Dongji, Jain, Vinija, Chadha, Aman, Sandage, Mary, Pope, Lauramarie, Dozier, Gerry, Seals, Cheryl |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
OffensiveLang: A Community Based Implicit Offensive Language Dataset
par: Das, Amit, et autres
Publié: (2024)
par: Das, Amit, et autres
Publié: (2024)
Investigating Hallucination in Conversations for Low Resource Languages
par: Das, Amit, et autres
Publié: (2025)
par: Das, Amit, et autres
Publié: (2025)
Towards Effective Authorship Attribution: Integrating Class-Incremental Learning
par: Rahgouy, Mostafa, et autres
Publié: (2024)
par: Rahgouy, Mostafa, et autres
Publié: (2024)
Are LLMs Ready to Replace Bangla Annotators?
par: Hasan, Md. Najib, et autres
Publié: (2026)
par: Hasan, Md. Najib, et autres
Publié: (2026)
Assessing LLM Reliability on Temporally Recent Open-Domain Questions
par: Krishnappa, Pushwitha, et autres
Publié: (2026)
par: Krishnappa, Pushwitha, et autres
Publié: (2026)
TRACEALIGN -- Tracing the Drift: Attributing Alignment Failures to Training-Time Belief Sources in LLMs
par: Das, Amitava, et autres
Publié: (2025)
par: Das, Amitava, et autres
Publié: (2025)
Guiding Vision-Language Model Selection for Visual Question-Answering Across Tasks, Domains, and Knowledge Types
par: Sinha, Neelabh, et autres
Publié: (2024)
par: Sinha, Neelabh, et autres
Publié: (2024)
Are Small Language Models Ready to Compete with Large Language Models for Practical Applications?
par: Sinha, Neelabh, et autres
Publié: (2024)
par: Sinha, Neelabh, et autres
Publié: (2024)
LLMsAgainstHate @ NLU of Devanagari Script Languages 2025: Hate Speech Detection and Target Identification in Devanagari Languages via Parameter Efficient Fine-Tuning of LLMs
par: Sidibomma, Rushendra, et autres
Publié: (2024)
par: Sidibomma, Rushendra, et autres
Publié: (2024)
Zero-Shot Multi-Label Classification of Bangla Documents: Large Decoders Vs. Classic Encoders
par: Sarkar, Souvika, et autres
Publié: (2025)
par: Sarkar, Souvika, et autres
Publié: (2025)
When Shallow Wins: Silent Failures and the Depth-Accuracy Paradox in Latent Reasoning
par: Sahoo, Subramanyam, et autres
Publié: (2026)
par: Sahoo, Subramanyam, et autres
Publié: (2026)
Reasoning or Rhetoric? An Empirical Analysis of Moral Reasoning Explanations in Large Language Models
par: Kasat, Aryan, et autres
Publié: (2026)
par: Kasat, Aryan, et autres
Publié: (2026)
Dial E for Ethical Enforcement: institutional VETO power as a governance primitive
par: Sahoo, Subramanyam, et autres
Publié: (2026)
par: Sahoo, Subramanyam, et autres
Publié: (2026)
The Reasoning Trap -- Logical Reasoning as a Mechanistic Pathway to Situational Awareness
par: Sahoo, Subramanyam, et autres
Publié: (2026)
par: Sahoo, Subramanyam, et autres
Publié: (2026)
Born With a Silver Spoon? Investigating Socioeconomic Bias in Large Language Models
par: Singh, Smriti, et autres
Publié: (2024)
par: Singh, Smriti, et autres
Publié: (2024)
AlignGuard-LoRA: Alignment-Preserving Fine-Tuning via Fisher-Guided Decomposition and Riemannian-Geodesic Collision Regularization
par: Das, Amitava, et autres
Publié: (2025)
par: Das, Amitava, et autres
Publié: (2025)
SAHOO: Safeguarded Alignment for High-Order Optimization Objectives in Recursive Self-Improvement
par: Sahoo, Subramanyam, et autres
Publié: (2026)
par: Sahoo, Subramanyam, et autres
Publié: (2026)
I Can't Believe It's Not Robust: Catastrophic Collapse of Safety Classifiers under Embedding Drift
par: Sahoo, Subramanyam, et autres
Publié: (2026)
par: Sahoo, Subramanyam, et autres
Publié: (2026)
Position: The Complexity of Perfect AI Alignment -- Formalizing the RLHF Trilemma
par: Sahoo, Subramanyam, et autres
Publié: (2025)
par: Sahoo, Subramanyam, et autres
Publié: (2025)
Exploring the Impact of Large Language Models on Recommender Systems: An Extensive Review
par: Vats, Arpita, et autres
Publié: (2024)
par: Vats, Arpita, et autres
Publié: (2024)
How Culturally Aware are Vision-Language Models?
par: Burda-Lassen, Olena, et autres
Publié: (2024)
par: Burda-Lassen, Olena, et autres
Publié: (2024)
Exploring the Frontier of Vision-Language Models: A Survey of Current Methodologies and Future Directions
par: Ghosh, Akash, et autres
Publié: (2024)
par: Ghosh, Akash, et autres
Publié: (2024)
D-STEER - Preference Alignment Techniques Learn to Behave, not to Believe -- Beneath the Surface, DPO as Steering Vector Perturbation in Activation Space
par: Raina, Samarth, et autres
Publié: (2025)
par: Raina, Samarth, et autres
Publié: (2025)
AlignMerge - Alignment-Preserving Large Language Model Merging via Fisher-Guided Geometric Constraints
par: Roy, Aniruddha, et autres
Publié: (2025)
par: Roy, Aniruddha, et autres
Publié: (2025)
A Comprehensive Survey of Accelerated Generation Techniques in Large Language Models
par: Khoshnoodi, Mahsa, et autres
Publié: (2024)
par: Khoshnoodi, Mahsa, et autres
Publié: (2024)
Neural FOXP2 -- Language Specific Neuron Steering for Targeted Language Improvement in LLMs
par: Saha, Anusa, et autres
Publié: (2026)
par: Saha, Anusa, et autres
Publié: (2026)
ECLIPTICA -- A Framework for Switchable LLM Alignment via CITA - Contrastive Instruction-Tuned Alignment
par: Wanaskar, Kapil, et autres
Publié: (2026)
par: Wanaskar, Kapil, et autres
Publié: (2026)
Personality Shapes Gender Bias in Persona-Conditioned LLM Narratives Across English and Hindi: An Empirical Investigation
par: Kumar, Tanay, et autres
Publié: (2026)
par: Kumar, Tanay, et autres
Publié: (2026)
Multilingual State Space Models for Structured Question Answering in Indic Languages
par: Vats, Arpita, et autres
Publié: (2025)
par: Vats, Arpita, et autres
Publié: (2025)
MOD-X: A Modular Open Decentralized eXchange Framework proposal for Heterogeneous Interoperable Artificial Intelligence Agents
par: Ioannides, Georgios, et autres
Publié: (2025)
par: Ioannides, Georgios, et autres
Publié: (2025)
Decoding the Diversity: A Review of the Indic AI Research Landscape
par: KJ, Sankalp, et autres
Publié: (2024)
par: KJ, Sankalp, et autres
Publié: (2024)
Refining Text-to-Image Generation: Towards Accurate Training-Free Glyph-Enhanced Image Generation
par: Lakhanpal, Sanyam, et autres
Publié: (2024)
par: Lakhanpal, Sanyam, et autres
Publié: (2024)
SKETCH: Structured Knowledge Enhanced Text Comprehension for Holistic Retrieval
par: Mahalingam, Aakash, et autres
Publié: (2024)
par: Mahalingam, Aakash, et autres
Publié: (2024)
LLM for Complex Reasoning Task: An Exploratory Study in Fermi Problems
par: Liu, Zishuo, et autres
Publié: (2025)
par: Liu, Zishuo, et autres
Publié: (2025)
DeHate: A Stable Diffusion-based Multimodal Approach to Mitigate Hate Speech in Images
par: Dalal, Dwip, et autres
Publié: (2025)
par: Dalal, Dwip, et autres
Publié: (2025)
Human-Readable Adversarial Prompts: An Investigation into LLM Vulnerabilities Using Situational Context
par: Das, Nilanjana, et autres
Publié: (2024)
par: Das, Nilanjana, et autres
Publié: (2024)
Parameter Efficient Fine Tuning: A Comprehensive Analysis Across Applications
par: Balne, Charith Chandra Sai, et autres
Publié: (2024)
par: Balne, Charith Chandra Sai, et autres
Publié: (2024)
On the Feasibility of Vision-Language Models for Time-Series Classification
par: Prithyani, Vinay, et autres
Publié: (2024)
par: Prithyani, Vinay, et autres
Publié: (2024)
Moral Sensitivity in LLMs: A Tiered Evaluation of Contextual Bias via Behavioral Profiling and Mechanistic Interpretability
par: Aggarwal, Yash, et autres
Publié: (2026)
par: Aggarwal, Yash, et autres
Publié: (2026)
Hierarchical Prompting Taxonomy: A Universal Evaluation Framework for Large Language Models Aligned with Human Cognitive Principles
par: Budagam, Devichand, et autres
Publié: (2024)
par: Budagam, Devichand, et autres
Publié: (2024)
Documents similaires
-
OffensiveLang: A Community Based Implicit Offensive Language Dataset
par: Das, Amit, et autres
Publié: (2024) -
Investigating Hallucination in Conversations for Low Resource Languages
par: Das, Amit, et autres
Publié: (2025) -
Towards Effective Authorship Attribution: Integrating Class-Incremental Learning
par: Rahgouy, Mostafa, et autres
Publié: (2024) -
Are LLMs Ready to Replace Bangla Annotators?
par: Hasan, Md. Najib, et autres
Publié: (2026) -
Assessing LLM Reliability on Temporally Recent Open-Domain Questions
par: Krishnappa, Pushwitha, et autres
Publié: (2026)