Exploring the Plausibility of Hate and Counter Speech Detectors with Explainable AI
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Böck, Adrian Jaques, Slijepčević, Djordje, Zeppelzauer, Matthias |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Distilling Knowledge from Large Language Models: A Concept Bottleneck Model for Hate and Counter Speech Recognition
von: Labadie-Tamayo, Roberto, et al.
Veröffentlicht: (2025)
von: Labadie-Tamayo, Roberto, et al.
Veröffentlicht: (2025)
FHSTP@EXIST 2025 Benchmark: Sexism Detection with Transparent Speech Concept Bottleneck Models
von: Labadie-Tamayo, Roberto, et al.
Veröffentlicht: (2025)
von: Labadie-Tamayo, Roberto, et al.
Veröffentlicht: (2025)
Explanatory Interactive Machine Learning for Bias Mitigation in Visual Gender Classification
von: Satriani, Nathanya, et al.
Veröffentlicht: (2026)
von: Satriani, Nathanya, et al.
Veröffentlicht: (2026)
Improving Hate Speech Classification with Cross-Taxonomy Dataset Integration
von: Fillies, Jan, et al.
Veröffentlicht: (2025)
von: Fillies, Jan, et al.
Veröffentlicht: (2025)
Towards Generalizable Generic Harmful Speech Datasets for Implicit Hate Speech Detection
von: Almohaimeed, Saad, et al.
Veröffentlicht: (2025)
von: Almohaimeed, Saad, et al.
Veröffentlicht: (2025)
Investigating Annotator Bias in Large Language Models for Hate Speech Detection
von: Das, Amit, et al.
Veröffentlicht: (2024)
von: Das, Amit, et al.
Veröffentlicht: (2024)
Causality Guided Representation Learning for Cross-Style Hate Speech Detection
von: Zhao, Chengshuai, et al.
Veröffentlicht: (2025)
von: Zhao, Chengshuai, et al.
Veröffentlicht: (2025)
Transformers and Ensemble methods: A solution for Hate Speech Detection in Arabic languages
von: de Paula, Angel Felipe Magnossão, et al.
Veröffentlicht: (2023)
von: de Paula, Angel Felipe Magnossão, et al.
Veröffentlicht: (2023)
Towards Interpretable Hate Speech Detection using Large Language Model-extracted Rationales
von: Nirmal, Ayushi, et al.
Veröffentlicht: (2024)
von: Nirmal, Ayushi, et al.
Veröffentlicht: (2024)
Explainable Identification of Hate Speech towards Islam using Graph Neural Networks
von: Wasi, Azmine Toushik
Veröffentlicht: (2023)
von: Wasi, Azmine Toushik
Veröffentlicht: (2023)
Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content
von: Stepanov, Ihor, et al.
Veröffentlicht: (2026)
von: Stepanov, Ihor, et al.
Veröffentlicht: (2026)
Exploring the Trade-off Between Model Performance and Explanation Plausibility of Text Classifiers Using Human Rationales
von: Resck, Lucas E., et al.
Veröffentlicht: (2024)
von: Resck, Lucas E., et al.
Veröffentlicht: (2024)
Multilingual Hate Speech Detection in Social Media Using Translation-Based Approaches with Large Language Models
von: Usman, Muhammad, et al.
Veröffentlicht: (2025)
von: Usman, Muhammad, et al.
Veröffentlicht: (2025)
Consolidating Strategies for Countering Hate Speech Using Persuasive Dialogues
von: Saha, Sougata, et al.
Veröffentlicht: (2024)
von: Saha, Sougata, et al.
Veröffentlicht: (2024)
Detecting Hope, Hate, and Emotion in Arabic Textual Speech and Multi-modal Memes Using Large Language Models
von: AlDahoul, Nouar, et al.
Veröffentlicht: (2025)
von: AlDahoul, Nouar, et al.
Veröffentlicht: (2025)
1-800-SHARED-TASKS @ NLU of Devanagari Script Languages: Detection of Language, Hate Speech, and Targets using LLMs
von: Purbey, Jebish, et al.
Veröffentlicht: (2024)
von: Purbey, Jebish, et al.
Veröffentlicht: (2024)
Parameter-Efficient Fine-Tuning for Low-Resource Languages: A Comparative Study of LLMs for Bengali Hate Speech Detection
von: Islam, Akif, et al.
Veröffentlicht: (2025)
von: Islam, Akif, et al.
Veröffentlicht: (2025)
BEExAI: Benchmark to Evaluate Explainable AI
von: Sithakoul, Samuel, et al.
Veröffentlicht: (2024)
von: Sithakoul, Samuel, et al.
Veröffentlicht: (2024)
Maximizing Mutual Information Between Prompt and Response Improves LLM Performance With No Additional Data
von: Nam, Hyunji, et al.
Veröffentlicht: (2026)
von: Nam, Hyunji, et al.
Veröffentlicht: (2026)
Base Models Look Human To AI Detectors
von: Xu, Yixuan Even, et al.
Veröffentlicht: (2026)
von: Xu, Yixuan Even, et al.
Veröffentlicht: (2026)
Think Like a Person Before Responding: A Multi-Faceted Evaluation of Persona-Guided LLMs for Countering Hate
von: Ngueajio, Mikel K., et al.
Veröffentlicht: (2025)
von: Ngueajio, Mikel K., et al.
Veröffentlicht: (2025)
Reframing Data Value for Large Language Models Through the Lens of Plausibility
von: Rammal, Mohamad Rida, et al.
Veröffentlicht: (2024)
von: Rammal, Mohamad Rida, et al.
Veröffentlicht: (2024)
OSPC: Artificial VLM Features for Hateful Meme Detection
von: Grönquist, Peter
Veröffentlicht: (2024)
von: Grönquist, Peter
Veröffentlicht: (2024)
Generating Plausible Distractors for Multiple-Choice Questions via Student Choice Prediction
von: Lee, Yooseop, et al.
Veröffentlicht: (2025)
von: Lee, Yooseop, et al.
Veröffentlicht: (2025)
Dealing with Annotator Disagreement in Hate Speech Classification
von: Dehghan, Somaiyeh, et al.
Veröffentlicht: (2025)
von: Dehghan, Somaiyeh, et al.
Veröffentlicht: (2025)
Navigating the Shadows: Unveiling Effective Disturbances for Modern AI Content Detectors
von: Zhou, Ying, et al.
Veröffentlicht: (2024)
von: Zhou, Ying, et al.
Veröffentlicht: (2024)
Hatred Stems from Ignorance! Distillation of the Persuasion Modes in Countering Conversational Hate Speech
von: Alyahya, Ghadi, et al.
Veröffentlicht: (2024)
von: Alyahya, Ghadi, et al.
Veröffentlicht: (2024)
An Investigation of Large Language Models for Real-World Hate Speech Detection
von: Guo, Keyan, et al.
Veröffentlicht: (2024)
von: Guo, Keyan, et al.
Veröffentlicht: (2024)
Infer Human's Intentions Before Following Natural Language Instructions
von: Wan, Yanming, et al.
Veröffentlicht: (2024)
von: Wan, Yanming, et al.
Veröffentlicht: (2024)
Amplifying, Not Learning: Fine-Tuned AI Text Detectors Amplify a Pretrained Direction
von: Smirnov, Alexander
Veröffentlicht: (2026)
von: Smirnov, Alexander
Veröffentlicht: (2026)
Self-Exploring Language Models for Explainable Link Forecasting on Temporal Graphs via Reinforcement Learning
von: Ding, Zifeng, et al.
Veröffentlicht: (2025)
von: Ding, Zifeng, et al.
Veröffentlicht: (2025)
IPAD: Inverse Prompt for AI Detection - A Robust and Interpretable LLM-Generated Text Detector
von: Chen, Zheng, et al.
Veröffentlicht: (2025)
von: Chen, Zheng, et al.
Veröffentlicht: (2025)
Evaluating & Reducing Deceptive Dialogue From Language Models with Multi-turn RL
von: Abdulhai, Marwa, et al.
Veröffentlicht: (2025)
von: Abdulhai, Marwa, et al.
Veröffentlicht: (2025)
Identifying False Content and Hate Speech in Sinhala YouTube Videos by Analyzing the Audio
von: Wickramaarachchi, W. A. K. M., et al.
Veröffentlicht: (2024)
von: Wickramaarachchi, W. A. K. M., et al.
Veröffentlicht: (2024)
RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization
von: Dong, Yihong, et al.
Veröffentlicht: (2025)
von: Dong, Yihong, et al.
Veröffentlicht: (2025)
Exploring and Improving Drafts in Blockwise Parallel Decoding
von: Kim, Taehyeon, et al.
Veröffentlicht: (2024)
von: Kim, Taehyeon, et al.
Veröffentlicht: (2024)
A Multilingual Sentiment Lexicon for Low-Resource Language Translation using Large Languages Models and Explainable AI
von: Malinga, Melusi, et al.
Veröffentlicht: (2024)
von: Malinga, Melusi, et al.
Veröffentlicht: (2024)
Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning
von: Poddar, Sriyash, et al.
Veröffentlicht: (2024)
von: Poddar, Sriyash, et al.
Veröffentlicht: (2024)
A Quantum Inspired Variational Kernel and Explainable AI Framework for Cross Region Solar and Wind Energy Forecasting
von: Manjunath, Pavan, et al.
Veröffentlicht: (2026)
von: Manjunath, Pavan, et al.
Veröffentlicht: (2026)
Inner Speech as Behavior Guides: Steerable Imitation of Diverse Behaviors for Human-AI coordination
von: Trivedi, Rakshit, et al.
Veröffentlicht: (2026)
von: Trivedi, Rakshit, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Distilling Knowledge from Large Language Models: A Concept Bottleneck Model for Hate and Counter Speech Recognition
von: Labadie-Tamayo, Roberto, et al.
Veröffentlicht: (2025) -
FHSTP@EXIST 2025 Benchmark: Sexism Detection with Transparent Speech Concept Bottleneck Models
von: Labadie-Tamayo, Roberto, et al.
Veröffentlicht: (2025) -
Explanatory Interactive Machine Learning for Bias Mitigation in Visual Gender Classification
von: Satriani, Nathanya, et al.
Veröffentlicht: (2026) -
Improving Hate Speech Classification with Cross-Taxonomy Dataset Integration
von: Fillies, Jan, et al.
Veröffentlicht: (2025) -
Towards Generalizable Generic Harmful Speech Datasets for Implicit Hate Speech Detection
von: Almohaimeed, Saad, et al.
Veröffentlicht: (2025)