PEACE 2.0: Grounded Explanations and Counter-Speech for Combating Hate Expressions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Damo, Greta, Petiot, Stéphane, Cabrio, Elena, Villata, Serena |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Effectiveness of Counter-Speech against Abusive Content: A Multidimensional Annotation and Classification Study
von: Damo, Greta, et al.
Veröffentlicht: (2025)
von: Damo, Greta, et al.
Veröffentlicht: (2025)
Beating Harmful Stereotypes Through Facts: RAG-based Counter-speech Generation
von: Damo, Greta, et al.
Veröffentlicht: (2025)
von: Damo, Greta, et al.
Veröffentlicht: (2025)
Is Safer Better? The Impact of Guardrails on the Argumentative Strength of LLMs in Hate Speech Countering
von: Bonaldi, Helena, et al.
Veröffentlicht: (2024)
von: Bonaldi, Helena, et al.
Veröffentlicht: (2024)
RooseBERT: A New Deal For Political Language Modelling
von: Dore, Deborah, et al.
Veröffentlicht: (2025)
von: Dore, Deborah, et al.
Veröffentlicht: (2025)
Compact Prompting in Instruction-tuned LLMs for Joint Argumentative Component Detection
von: Elguendouze, Sofiane, et al.
Veröffentlicht: (2026)
von: Elguendouze, Sofiane, et al.
Veröffentlicht: (2026)
CasiMedicos-Arg: A Medical Question Answering Dataset Annotated with Explanatory Argumentative Structures
von: Sviridova, Ekaterina, et al.
Veröffentlicht: (2024)
von: Sviridova, Ekaterina, et al.
Veröffentlicht: (2024)
Argument Quality Assessment in the Age of Instruction-Following Large Language Models
von: Wachsmuth, Henning, et al.
Veröffentlicht: (2024)
von: Wachsmuth, Henning, et al.
Veröffentlicht: (2024)
HateGPT: Unleashing GPT-3.5 Turbo to Combat Hate Speech on X
von: Deroy, Aniket, et al.
Veröffentlicht: (2024)
von: Deroy, Aniket, et al.
Veröffentlicht: (2024)
Outcome-Constrained Large Language Models for Countering Hate Speech
von: Hong, Lingzi, et al.
Veröffentlicht: (2024)
von: Hong, Lingzi, et al.
Veröffentlicht: (2024)
Dialogues of Dissent: Thematic and Rhetorical Dimensions of Hate and Counter-Hate Speech in Social Media Conversations
von: Levi, Effi, et al.
Veröffentlicht: (2025)
von: Levi, Effi, et al.
Veröffentlicht: (2025)
Consolidating Strategies for Countering Hate Speech Using Persuasive Dialogues
von: Saha, Sougata, et al.
Veröffentlicht: (2024)
von: Saha, Sougata, et al.
Veröffentlicht: (2024)
Incorporating Human Explanations for Robust Hate Speech Detection
von: Chen, Jennifer L., et al.
Veröffentlicht: (2024)
von: Chen, Jennifer L., et al.
Veröffentlicht: (2024)
Exploring the Plausibility of Hate and Counter Speech Detectors with Explainable AI
von: Böck, Adrian Jaques, et al.
Veröffentlicht: (2024)
von: Böck, Adrian Jaques, et al.
Veröffentlicht: (2024)
ReZG: Retrieval-Augmented Zero-Shot Counter Narrative Generation for Hate Speech
von: Jiang, Shuyu, et al.
Veröffentlicht: (2023)
von: Jiang, Shuyu, et al.
Veröffentlicht: (2023)
HateXScore: A Metric Suite for Evaluating Reasoning Quality in Hate Speech Explanations
von: Hu, Yujia, et al.
Veröffentlicht: (2026)
von: Hu, Yujia, et al.
Veröffentlicht: (2026)
WATCHED: A Web AI Agent Tool for Combating Hate Speech by Expanding Data
von: Piot, Paloma, et al.
Veröffentlicht: (2025)
von: Piot, Paloma, et al.
Veröffentlicht: (2025)
Harnessing Artificial Intelligence to Combat Online Hate: Exploring the Challenges and Opportunities of Large Language Models in Hate Speech Detection
von: Kumarage, Tharindu, et al.
Veröffentlicht: (2024)
von: Kumarage, Tharindu, et al.
Veröffentlicht: (2024)
Hatred Stems from Ignorance! Distillation of the Persuasion Modes in Countering Conversational Hate Speech
von: Alyahya, Ghadi, et al.
Veröffentlicht: (2024)
von: Alyahya, Ghadi, et al.
Veröffentlicht: (2024)
Tox-BART: Leveraging Toxicity Attributes for Explanation Generation of Implicit Hate Speech
von: Yadav, Neemesh, et al.
Veröffentlicht: (2024)
von: Yadav, Neemesh, et al.
Veröffentlicht: (2024)
Explainability and Hate Speech: Structured Explanations Make Social Media Moderators Faster
von: Calabrese, Agostina, et al.
Veröffentlicht: (2024)
von: Calabrese, Agostina, et al.
Veröffentlicht: (2024)
ChildGuard: A Specialized Dataset for Combatting Child-Targeted Hate Speech
von: Kashyap, Gautam Siddharth, et al.
Veröffentlicht: (2025)
von: Kashyap, Gautam Siddharth, et al.
Veröffentlicht: (2025)
HateDebias: On the Diversity and Variability of Hate Speech Debiasing
von: Wu, Hongyan, et al.
Veröffentlicht: (2024)
von: Wu, Hongyan, et al.
Veröffentlicht: (2024)
Decoding Hate: Exploring Language Models' Reactions to Hate Speech
von: Piot, Paloma, et al.
Veröffentlicht: (2024)
von: Piot, Paloma, et al.
Veröffentlicht: (2024)
COT: A Generative Approach for Hate Speech Counter-Narratives via Contrastive Optimal Transport
von: Zhang, Linhao, et al.
Veröffentlicht: (2024)
von: Zhang, Linhao, et al.
Veröffentlicht: (2024)
Bridging Fairness and Explainability: Can Input-Based Explanations Promote Fairness in Hate Speech Detection?
von: Wang, Yifan, et al.
Veröffentlicht: (2025)
von: Wang, Yifan, et al.
Veröffentlicht: (2025)
"Is Hate Lost in Translation?": Evaluation of Multilingual LGBTQIA+ Hate Speech Detection
von: Chan, Fai Leui, et al.
Veröffentlicht: (2024)
von: Chan, Fai Leui, et al.
Veröffentlicht: (2024)
Web(er) of Hate: A Survey on How Hate Speech Is Typed
von: Wang, Luna, et al.
Veröffentlicht: (2025)
von: Wang, Luna, et al.
Veröffentlicht: (2025)
Hateful Person or Hateful Model? Investigating the Role of Personas in Hate Speech Detection by Large Language Models
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2025)
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2025)
LLM in the Loop: Creating the ParaDeHate Dataset for Hate Speech Detoxification
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2025)
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2025)
MFTCXplain: A Multilingual Benchmark Dataset for Evaluating the Moral Reasoning of LLMs through Multi-hop Hate Speech Explanation
von: Trager, Jackson, et al.
Veröffentlicht: (2025)
von: Trager, Jackson, et al.
Veröffentlicht: (2025)
Stakeholder Suite: A Unified AI Framework for Mapping Actors, Topics and Arguments in Public Debates
von: Chenene, Mohamed, et al.
Veröffentlicht: (2025)
von: Chenene, Mohamed, et al.
Veröffentlicht: (2025)
MasonPerplexity at Multimodal Hate Speech Event Detection 2024: Hate Speech and Target Detection Using Transformer Ensembles
von: Ganguly, Amrita, et al.
Veröffentlicht: (2024)
von: Ganguly, Amrita, et al.
Veröffentlicht: (2024)
When Hate Meets Facts: LLMs-in-the-Loop for Check-worthiness Detection in Hate Speech
von: Ocampo, Nicolás Benjamín, et al.
Veröffentlicht: (2026)
von: Ocampo, Nicolás Benjamín, et al.
Veröffentlicht: (2026)
NaijaHate: Evaluating Hate Speech Detection on Nigerian Twitter Using Representative Data
von: Tonneau, Manuel, et al.
Veröffentlicht: (2024)
von: Tonneau, Manuel, et al.
Veröffentlicht: (2024)
Challenger at MultiPRIDE: Is It Hate Speech or Reclaimed?
von: Tekanlou, Hadi Bayrami Asl, et al.
Veröffentlicht: (2026)
von: Tekanlou, Hadi Bayrami Asl, et al.
Veröffentlicht: (2026)
Compositional Generalisation for Explainable Hate Speech Detection
von: Calabrese, Agostina, et al.
Veröffentlicht: (2025)
von: Calabrese, Agostina, et al.
Veröffentlicht: (2025)
Automatic Textual Normalization for Hate Speech Detection
von: Nguyen, Anh Thi-Hoang, et al.
Veröffentlicht: (2023)
von: Nguyen, Anh Thi-Hoang, et al.
Veröffentlicht: (2023)
Advancing Hate Speech Detection with Transformers: Insights from the MetaHate
von: Chapagain, Santosh, et al.
Veröffentlicht: (2025)
von: Chapagain, Santosh, et al.
Veröffentlicht: (2025)
Towards Effective Counter-Responses: Aligning Human Preferences with Strategies to Combat Online Trolling
von: Lee, Huije, et al.
Veröffentlicht: (2024)
von: Lee, Huije, et al.
Veröffentlicht: (2024)
The Enforcement and Feasibility of Hate Speech Moderation on Twitter
von: Tonneau, Manuel, et al.
Veröffentlicht: (2026)
von: Tonneau, Manuel, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Effectiveness of Counter-Speech against Abusive Content: A Multidimensional Annotation and Classification Study
von: Damo, Greta, et al.
Veröffentlicht: (2025) -
Beating Harmful Stereotypes Through Facts: RAG-based Counter-speech Generation
von: Damo, Greta, et al.
Veröffentlicht: (2025) -
Is Safer Better? The Impact of Guardrails on the Argumentative Strength of LLMs in Hate Speech Countering
von: Bonaldi, Helena, et al.
Veröffentlicht: (2024) -
RooseBERT: A New Deal For Political Language Modelling
von: Dore, Deborah, et al.
Veröffentlicht: (2025) -
Compact Prompting in Instruction-tuned LLMs for Joint Argumentative Component Detection
von: Elguendouze, Sofiane, et al.
Veröffentlicht: (2026)