Don't Go To Extremes: Revealing the Excessive Sensitivity and Calibration Limitations of LLMs in Implicit Hate Speech Detection
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Min, He, Jianfeng, Ji, Taoran, Lu, Chang-Tien |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Leveraging LLMs for Context-Aware Implicit Textual and Multimodal Hate Speech Detection
di: Brook, Joshua Wolfe, et al.
Pubblicazione: (2025)
di: Brook, Joshua Wolfe, et al.
Pubblicazione: (2025)
Cracking the Code: Enhancing Implicit Hate Speech Detection through Coding Classification
di: Wei, Lu, et al.
Pubblicazione: (2025)
di: Wei, Lu, et al.
Pubblicazione: (2025)
Show, Don't Tell: Uncovering Implicit Character Portrayal using LLMs
di: Jaipersaud, Brandon, et al.
Pubblicazione: (2024)
di: Jaipersaud, Brandon, et al.
Pubblicazione: (2024)
HatePrototypes: Interpretable and Transferable Representations for Implicit and Explicit Hate Speech Detection
di: Proskurina, Irina, et al.
Pubblicazione: (2025)
di: Proskurina, Irina, et al.
Pubblicazione: (2025)
Selective Demonstration Retrieval for Improved Implicit Hate Speech Detection
di: Kim, Yumin, et al.
Pubblicazione: (2025)
di: Kim, Yumin, et al.
Pubblicazione: (2025)
AMA-LSTM: Pioneering Robust and Fair Financial Audio Analysis for Stock Volatility Prediction
di: Wang, Shengkun, et al.
Pubblicazione: (2024)
di: Wang, Shengkun, et al.
Pubblicazione: (2024)
When Hate Meets Facts: LLMs-in-the-Loop for Check-worthiness Detection in Hate Speech
di: Ocampo, Nicolás Benjamín, et al.
Pubblicazione: (2026)
di: Ocampo, Nicolás Benjamín, et al.
Pubblicazione: (2026)
GPT-HateCheck: Can LLMs Write Better Functional Tests for Hate Speech Detection?
di: Jin, Yiping, et al.
Pubblicazione: (2024)
di: Jin, Yiping, et al.
Pubblicazione: (2024)
Towards Generalizable Generic Harmful Speech Datasets for Implicit Hate Speech Detection
di: Almohaimeed, Saad, et al.
Pubblicazione: (2025)
di: Almohaimeed, Saad, et al.
Pubblicazione: (2025)
Specializing General-purpose LLM Embeddings for Implicit Hate Speech Detection across Datasets
di: Cheremetiev, Vassiliy, et al.
Pubblicazione: (2025)
di: Cheremetiev, Vassiliy, et al.
Pubblicazione: (2025)
Don't Think Twice! Over-Reasoning Impairs Confidence Calibration
di: Lacombe, Romain, et al.
Pubblicazione: (2025)
di: Lacombe, Romain, et al.
Pubblicazione: (2025)
Label-aware Hard Negative Sampling Strategies with Momentum Contrastive Learning for Implicit Hate Speech Detection
di: Kim, Jaehoon, et al.
Pubblicazione: (2024)
di: Kim, Jaehoon, et al.
Pubblicazione: (2024)
"Is Hate Lost in Translation?": Evaluation of Multilingual LGBTQIA+ Hate Speech Detection
di: Chan, Fai Leui, et al.
Pubblicazione: (2024)
di: Chan, Fai Leui, et al.
Pubblicazione: (2024)
Tox-BART: Leveraging Toxicity Attributes for Explanation Generation of Implicit Hate Speech
di: Yadav, Neemesh, et al.
Pubblicazione: (2024)
di: Yadav, Neemesh, et al.
Pubblicazione: (2024)
Improving Implicit Hate Speech Detection via a Community-Driven Multi-Agent Framework
di: Gajewska, Ewelina, et al.
Pubblicazione: (2026)
di: Gajewska, Ewelina, et al.
Pubblicazione: (2026)
MasonPerplexity at Multimodal Hate Speech Event Detection 2024: Hate Speech and Target Detection Using Transformer Ensembles
di: Ganguly, Amrita, et al.
Pubblicazione: (2024)
di: Ganguly, Amrita, et al.
Pubblicazione: (2024)
Hateful Person or Hateful Model? Investigating the Role of Personas in Hate Speech Detection by Large Language Models
di: Yuan, Shuzhou, et al.
Pubblicazione: (2025)
di: Yuan, Shuzhou, et al.
Pubblicazione: (2025)
Algorithmic Fairness in NLP: Persona-Infused LLMs for Human-Centric Hate Speech Detection
di: Gajewska, Ewelina, et al.
Pubblicazione: (2025)
di: Gajewska, Ewelina, et al.
Pubblicazione: (2025)
Rethinking Hate Speech Detection on Social Media: Can LLMs Replace Traditional Models?
di: Singh, Daman Deep, et al.
Pubblicazione: (2025)
di: Singh, Daman Deep, et al.
Pubblicazione: (2025)
Calibrate, Don't Curate: Label-Efficient Estimation from Noisy LLM Judges
di: Li, Yanran
Pubblicazione: (2026)
di: Li, Yanran
Pubblicazione: (2026)
Compositional Generalisation for Explainable Hate Speech Detection
di: Calabrese, Agostina, et al.
Pubblicazione: (2025)
di: Calabrese, Agostina, et al.
Pubblicazione: (2025)
Automatic Textual Normalization for Hate Speech Detection
di: Nguyen, Anh Thi-Hoang, et al.
Pubblicazione: (2023)
di: Nguyen, Anh Thi-Hoang, et al.
Pubblicazione: (2023)
AmpleHate: Amplifying the Attention for Versatile Implicit Hate Detection
di: Lee, Yejin, et al.
Pubblicazione: (2025)
di: Lee, Yejin, et al.
Pubblicazione: (2025)
Be like a Goldfish, Don't Memorize! Mitigating Memorization in Generative LLMs
di: Hans, Abhimanyu, et al.
Pubblicazione: (2024)
di: Hans, Abhimanyu, et al.
Pubblicazione: (2024)
NaijaHate: Evaluating Hate Speech Detection on Nigerian Twitter Using Representative Data
di: Tonneau, Manuel, et al.
Pubblicazione: (2024)
di: Tonneau, Manuel, et al.
Pubblicazione: (2024)
Advancing Hate Speech Detection with Transformers: Insights from the MetaHate
di: Chapagain, Santosh, et al.
Pubblicazione: (2025)
di: Chapagain, Santosh, et al.
Pubblicazione: (2025)
RV-HATE: Reinforced Multi-Module Voting for Implicit Hate Speech Detection
di: Lee, Yejin, et al.
Pubblicazione: (2025)
di: Lee, Yejin, et al.
Pubblicazione: (2025)
LLMsAgainstHate @ NLU of Devanagari Script Languages 2025: Hate Speech Detection and Target Identification in Devanagari Languages via Parameter Efficient Fine-Tuning of LLMs
di: Sidibomma, Rushendra, et al.
Pubblicazione: (2024)
di: Sidibomma, Rushendra, et al.
Pubblicazione: (2024)
Don't Touch My Diacritics
di: Gorman, Kyle, et al.
Pubblicazione: (2024)
di: Gorman, Kyle, et al.
Pubblicazione: (2024)
Don't Pay Attention
di: Hammoud, Mohammad, et al.
Pubblicazione: (2025)
di: Hammoud, Mohammad, et al.
Pubblicazione: (2025)
Exploring the Deceptive Power of LLM-Generated Fake News: A Study of Real-World Detection Challenges
di: Sun, Yanshen, et al.
Pubblicazione: (2024)
di: Sun, Yanshen, et al.
Pubblicazione: (2024)
Self-Explaining Hate Speech Detection with Moral Rationales
di: Vargas, Francielle, et al.
Pubblicazione: (2026)
di: Vargas, Francielle, et al.
Pubblicazione: (2026)
Hate Speech According to the Law: An Analysis for Effective Detection
di: Korre, Katerina, et al.
Pubblicazione: (2024)
di: Korre, Katerina, et al.
Pubblicazione: (2024)
Incorporating Human Explanations for Robust Hate Speech Detection
di: Chen, Jennifer L., et al.
Pubblicazione: (2024)
di: Chen, Jennifer L., et al.
Pubblicazione: (2024)
Code-Mixed Telugu-English Hate Speech Detection
di: Kakarla, Santhosh, et al.
Pubblicazione: (2025)
di: Kakarla, Santhosh, et al.
Pubblicazione: (2025)
Hatevolution: What Static Benchmarks Don't Tell Us
di: Di Bonaventura, Chiara, et al.
Pubblicazione: (2025)
di: Di Bonaventura, Chiara, et al.
Pubblicazione: (2025)
Multi3Hate: Multimodal, Multilingual, and Multicultural Hate Speech Detection with Vision-Language Models
di: Bui, Minh Duc, et al.
Pubblicazione: (2024)
di: Bui, Minh Duc, et al.
Pubblicazione: (2024)
Don't Forget Your Reward Values: Language Model Alignment via Value-based Calibration
di: Mao, Xin, et al.
Pubblicazione: (2024)
di: Mao, Xin, et al.
Pubblicazione: (2024)
MetaHate: A Dataset for Unifying Efforts on Hate Speech Detection
di: Piot, Paloma, et al.
Pubblicazione: (2024)
di: Piot, Paloma, et al.
Pubblicazione: (2024)
HateDebias: On the Diversity and Variability of Hate Speech Debiasing
di: Wu, Hongyan, et al.
Pubblicazione: (2024)
di: Wu, Hongyan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Leveraging LLMs for Context-Aware Implicit Textual and Multimodal Hate Speech Detection
di: Brook, Joshua Wolfe, et al.
Pubblicazione: (2025) -
Cracking the Code: Enhancing Implicit Hate Speech Detection through Coding Classification
di: Wei, Lu, et al.
Pubblicazione: (2025) -
Show, Don't Tell: Uncovering Implicit Character Portrayal using LLMs
di: Jaipersaud, Brandon, et al.
Pubblicazione: (2024) -
HatePrototypes: Interpretable and Transferable Representations for Implicit and Explicit Hate Speech Detection
di: Proskurina, Irina, et al.
Pubblicazione: (2025) -
Selective Demonstration Retrieval for Improved Implicit Hate Speech Detection
di: Kim, Yumin, et al.
Pubblicazione: (2025)