Is Safer Better? The Impact of Guardrails on the Argumentative Strength of LLMs in Hate Speech Countering
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bonaldi, Helena, Damo, Greta, Ocampo, Nicolás Benjamín, Cabrio, Elena, Villata, Serena, Guerini, Marco |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PEACE 2.0: Grounded Explanations and Counter-Speech for Combating Hate Expressions
von: Damo, Greta, et al.
Veröffentlicht: (2026)
von: Damo, Greta, et al.
Veröffentlicht: (2026)
Effectiveness of Counter-Speech against Abusive Content: A Multidimensional Annotation and Classification Study
von: Damo, Greta, et al.
Veröffentlicht: (2025)
von: Damo, Greta, et al.
Veröffentlicht: (2025)
Beating Harmful Stereotypes Through Facts: RAG-based Counter-speech Generation
von: Damo, Greta, et al.
Veröffentlicht: (2025)
von: Damo, Greta, et al.
Veröffentlicht: (2025)
Assisted Counterspeech Writing at the Crossroads of Hate Speech and Misinformation
von: Martone, Genoveffa, et al.
Veröffentlicht: (2026)
von: Martone, Genoveffa, et al.
Veröffentlicht: (2026)
Compact Prompting in Instruction-tuned LLMs for Joint Argumentative Component Detection
von: Elguendouze, Sofiane, et al.
Veröffentlicht: (2026)
von: Elguendouze, Sofiane, et al.
Veröffentlicht: (2026)
RooseBERT: A New Deal For Political Language Modelling
von: Dore, Deborah, et al.
Veröffentlicht: (2025)
von: Dore, Deborah, et al.
Veröffentlicht: (2025)
NLP for Counterspeech against Hate: A Survey and How-To Guide
von: Bonaldi, Helena, et al.
Veröffentlicht: (2024)
von: Bonaldi, Helena, et al.
Veröffentlicht: (2024)
CasiMedicos-Arg: A Medical Question Answering Dataset Annotated with Explanatory Argumentative Structures
von: Sviridova, Ekaterina, et al.
Veröffentlicht: (2024)
von: Sviridova, Ekaterina, et al.
Veröffentlicht: (2024)
Leveraging Argument Structure to Predict Content Hatefulness
von: Ocampo, Nicolás Benjamín, et al.
Veröffentlicht: (2026)
von: Ocampo, Nicolás Benjamín, et al.
Veröffentlicht: (2026)
Argument Quality Assessment in the Age of Instruction-Following Large Language Models
von: Wachsmuth, Henning, et al.
Veröffentlicht: (2024)
von: Wachsmuth, Henning, et al.
Veröffentlicht: (2024)
When Hate Meets Facts: LLMs-in-the-Loop for Check-worthiness Detection in Hate Speech
von: Ocampo, Nicolás Benjamín, et al.
Veröffentlicht: (2026)
von: Ocampo, Nicolás Benjamín, et al.
Veröffentlicht: (2026)
GPT-HateCheck: Can LLMs Write Better Functional Tests for Hate Speech Detection?
von: Jin, Yiping, et al.
Veröffentlicht: (2024)
von: Jin, Yiping, et al.
Veröffentlicht: (2024)
Outcome-Constrained Large Language Models for Countering Hate Speech
von: Hong, Lingzi, et al.
Veröffentlicht: (2024)
von: Hong, Lingzi, et al.
Veröffentlicht: (2024)
Stakeholder Suite: A Unified AI Framework for Mapping Actors, Topics and Arguments in Public Debates
von: Chenene, Mohamed, et al.
Veröffentlicht: (2025)
von: Chenene, Mohamed, et al.
Veröffentlicht: (2025)
Dialogues of Dissent: Thematic and Rhetorical Dimensions of Hate and Counter-Hate Speech in Social Media Conversations
von: Levi, Effi, et al.
Veröffentlicht: (2025)
von: Levi, Effi, et al.
Veröffentlicht: (2025)
Consolidating Strategies for Countering Hate Speech Using Persuasive Dialogues
von: Saha, Sougata, et al.
Veröffentlicht: (2024)
von: Saha, Sougata, et al.
Veröffentlicht: (2024)
Basque and Spanish Counter Narrative Generation: Data Creation and Evaluation
von: Bengoetxea, Jaione, et al.
Veröffentlicht: (2024)
von: Bengoetxea, Jaione, et al.
Veröffentlicht: (2024)
Exploring the Plausibility of Hate and Counter Speech Detectors with Explainable AI
von: Böck, Adrian Jaques, et al.
Veröffentlicht: (2024)
von: Böck, Adrian Jaques, et al.
Veröffentlicht: (2024)
ReZG: Retrieval-Augmented Zero-Shot Counter Narrative Generation for Hate Speech
von: Jiang, Shuyu, et al.
Veröffentlicht: (2023)
von: Jiang, Shuyu, et al.
Veröffentlicht: (2023)
Argument Quality Assessment with Large Language Models: A Pairwise Bradley-Terry Approach
von: Ocampo, Nicolás Benjamín, et al.
Veröffentlicht: (2026)
von: Ocampo, Nicolás Benjamín, et al.
Veröffentlicht: (2026)
Hatred Stems from Ignorance! Distillation of the Persuasion Modes in Countering Conversational Hate Speech
von: Alyahya, Ghadi, et al.
Veröffentlicht: (2024)
von: Alyahya, Ghadi, et al.
Veröffentlicht: (2024)
Alternative Speech: Complementary Method to Counter-Narrative for Better Discourse
von: Lee, Seungyoon, et al.
Veröffentlicht: (2024)
von: Lee, Seungyoon, et al.
Veröffentlicht: (2024)
Guardrail Baselines for Unlearning in LLMs
von: Thaker, Pratiksha, et al.
Veröffentlicht: (2024)
von: Thaker, Pratiksha, et al.
Veröffentlicht: (2024)
HateDebias: On the Diversity and Variability of Hate Speech Debiasing
von: Wu, Hongyan, et al.
Veröffentlicht: (2024)
von: Wu, Hongyan, et al.
Veröffentlicht: (2024)
Stronger Together: Unleashing the Social Impact of Hate Speech Research
von: Wong, Sidney
Veröffentlicht: (2025)
von: Wong, Sidney
Veröffentlicht: (2025)
Decoding Hate: Exploring Language Models' Reactions to Hate Speech
von: Piot, Paloma, et al.
Veröffentlicht: (2024)
von: Piot, Paloma, et al.
Veröffentlicht: (2024)
Leveraging LLMs for Context-Aware Implicit Textual and Multimodal Hate Speech Detection
von: Brook, Joshua Wolfe, et al.
Veröffentlicht: (2025)
von: Brook, Joshua Wolfe, et al.
Veröffentlicht: (2025)
COT: A Generative Approach for Hate Speech Counter-Narratives via Contrastive Optimal Transport
von: Zhang, Linhao, et al.
Veröffentlicht: (2024)
von: Zhang, Linhao, et al.
Veröffentlicht: (2024)
"Is Hate Lost in Translation?": Evaluation of Multilingual LGBTQIA+ Hate Speech Detection
von: Chan, Fai Leui, et al.
Veröffentlicht: (2024)
von: Chan, Fai Leui, et al.
Veröffentlicht: (2024)
Web(er) of Hate: A Survey on How Hate Speech Is Typed
von: Wang, Luna, et al.
Veröffentlicht: (2025)
von: Wang, Luna, et al.
Veröffentlicht: (2025)
Hateful Person or Hateful Model? Investigating the Role of Personas in Hate Speech Detection by Large Language Models
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2025)
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2025)
"Reasoning" with Rhetoric: On the Style-Evidence Tradeoff in LLM-Generated Counter-Arguments
von: Verma, Preetika, et al.
Veröffentlicht: (2024)
von: Verma, Preetika, et al.
Veröffentlicht: (2024)
LLMsAgainstHate @ NLU of Devanagari Script Languages 2025: Hate Speech Detection and Target Identification in Devanagari Languages via Parameter Efficient Fine-Tuning of LLMs
von: Sidibomma, Rushendra, et al.
Veröffentlicht: (2024)
von: Sidibomma, Rushendra, et al.
Veröffentlicht: (2024)
Algorithmic Fairness in NLP: Persona-Infused LLMs for Human-Centric Hate Speech Detection
von: Gajewska, Ewelina, et al.
Veröffentlicht: (2025)
von: Gajewska, Ewelina, et al.
Veröffentlicht: (2025)
Rethinking Hate Speech Detection on Social Media: Can LLMs Replace Traditional Models?
von: Singh, Daman Deep, et al.
Veröffentlicht: (2025)
von: Singh, Daman Deep, et al.
Veröffentlicht: (2025)
LLM in the Loop: Creating the ParaDeHate Dataset for Hate Speech Detoxification
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2025)
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2025)
MasonPerplexity at Multimodal Hate Speech Event Detection 2024: Hate Speech and Target Detection Using Transformer Ensembles
von: Ganguly, Amrita, et al.
Veröffentlicht: (2024)
von: Ganguly, Amrita, et al.
Veröffentlicht: (2024)
HateGPT: Unleashing GPT-3.5 Turbo to Combat Hate Speech on X
von: Deroy, Aniket, et al.
Veröffentlicht: (2024)
von: Deroy, Aniket, et al.
Veröffentlicht: (2024)
NaijaHate: Evaluating Hate Speech Detection on Nigerian Twitter Using Representative Data
von: Tonneau, Manuel, et al.
Veröffentlicht: (2024)
von: Tonneau, Manuel, et al.
Veröffentlicht: (2024)
Compositional Generalisation for Explainable Hate Speech Detection
von: Calabrese, Agostina, et al.
Veröffentlicht: (2025)
von: Calabrese, Agostina, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
PEACE 2.0: Grounded Explanations and Counter-Speech for Combating Hate Expressions
von: Damo, Greta, et al.
Veröffentlicht: (2026) -
Effectiveness of Counter-Speech against Abusive Content: A Multidimensional Annotation and Classification Study
von: Damo, Greta, et al.
Veröffentlicht: (2025) -
Beating Harmful Stereotypes Through Facts: RAG-based Counter-speech Generation
von: Damo, Greta, et al.
Veröffentlicht: (2025) -
Assisted Counterspeech Writing at the Crossroads of Hate Speech and Misinformation
von: Martone, Genoveffa, et al.
Veröffentlicht: (2026) -
Compact Prompting in Instruction-tuned LLMs for Joint Argumentative Component Detection
von: Elguendouze, Sofiane, et al.
Veröffentlicht: (2026)