Towards Interpretable Hate Speech Detection using Large Language Model-extracted Rationales
Fuente:
arXiv
Salvato in:
| Autori principali: | Nirmal, Ayushi, Bhattacharjee, Amrita, Sheth, Paras, Liu, Huan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Causality Guided Representation Learning for Cross-Style Hate Speech Detection
di: Zhao, Chengshuai, et al.
Pubblicazione: (2025)
di: Zhao, Chengshuai, et al.
Pubblicazione: (2025)
Harnessing Artificial Intelligence to Combat Online Hate: Exploring the Challenges and Opportunities of Large Language Models in Hate Speech Detection
di: Kumarage, Tharindu, et al.
Pubblicazione: (2024)
di: Kumarage, Tharindu, et al.
Pubblicazione: (2024)
Cross-Platform Hate Speech Detection with Weakly Supervised Causal Disentanglement
di: Sheth, Paras, et al.
Pubblicazione: (2024)
di: Sheth, Paras, et al.
Pubblicazione: (2024)
Adversarial Text Purification: A Large Language Model Approach for Defense
di: Moraffah, Raha, et al.
Pubblicazione: (2024)
di: Moraffah, Raha, et al.
Pubblicazione: (2024)
Towards LLM-guided Causal Explainability for Black-box Text Classifiers
di: Bhattacharjee, Amrita, et al.
Pubblicazione: (2023)
di: Bhattacharjee, Amrita, et al.
Pubblicazione: (2023)
EAGLE: A Domain Generalization Framework for AI-generated Text Detection
di: Bhattacharjee, Amrita, et al.
Pubblicazione: (2024)
di: Bhattacharjee, Amrita, et al.
Pubblicazione: (2024)
Investigating Annotator Bias in Large Language Models for Hate Speech Detection
di: Das, Amit, et al.
Pubblicazione: (2024)
di: Das, Amit, et al.
Pubblicazione: (2024)
Zero-shot LLM-guided Counterfactual Generation: A Case Study on NLP Model Evaluation
di: Bhattacharjee, Amrita, et al.
Pubblicazione: (2024)
di: Bhattacharjee, Amrita, et al.
Pubblicazione: (2024)
Towards Generalizable Generic Harmful Speech Datasets for Implicit Hate Speech Detection
di: Almohaimeed, Saad, et al.
Pubblicazione: (2025)
di: Almohaimeed, Saad, et al.
Pubblicazione: (2025)
Multilingual Hate Speech Detection in Social Media Using Translation-Based Approaches with Large Language Models
di: Usman, Muhammad, et al.
Pubblicazione: (2025)
di: Usman, Muhammad, et al.
Pubblicazione: (2025)
Detecting Hope, Hate, and Emotion in Arabic Textual Speech and Multi-modal Memes Using Large Language Models
di: AlDahoul, Nouar, et al.
Pubblicazione: (2025)
di: AlDahoul, Nouar, et al.
Pubblicazione: (2025)
An Investigation of Large Language Models for Real-World Hate Speech Detection
di: Guo, Keyan, et al.
Pubblicazione: (2024)
di: Guo, Keyan, et al.
Pubblicazione: (2024)
1-800-SHARED-TASKS @ NLU of Devanagari Script Languages: Detection of Language, Hate Speech, and Targets using LLMs
di: Purbey, Jebish, et al.
Pubblicazione: (2024)
di: Purbey, Jebish, et al.
Pubblicazione: (2024)
Transformers and Ensemble methods: A solution for Hate Speech Detection in Arabic languages
di: de Paula, Angel Felipe Magnossão, et al.
Pubblicazione: (2023)
di: de Paula, Angel Felipe Magnossão, et al.
Pubblicazione: (2023)
Neural Diversity Regularizes Hallucinations in Language Models
di: Chakrabarti, Kushal, et al.
Pubblicazione: (2025)
di: Chakrabarti, Kushal, et al.
Pubblicazione: (2025)
Parameter-Efficient Fine-Tuning for Low-Resource Languages: A Comparative Study of LLMs for Bengali Hate Speech Detection
di: Islam, Akif, et al.
Pubblicazione: (2025)
di: Islam, Akif, et al.
Pubblicazione: (2025)
Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
di: Lan, Michael, et al.
Pubblicazione: (2023)
di: Lan, Michael, et al.
Pubblicazione: (2023)
Exploring the Plausibility of Hate and Counter Speech Detectors with Explainable AI
di: Böck, Adrian Jaques, et al.
Pubblicazione: (2024)
di: Böck, Adrian Jaques, et al.
Pubblicazione: (2024)
Distilling Knowledge from Large Language Models: A Concept Bottleneck Model for Hate and Counter Speech Recognition
di: Labadie-Tamayo, Roberto, et al.
Pubblicazione: (2025)
di: Labadie-Tamayo, Roberto, et al.
Pubblicazione: (2025)
Mechanistic Interpretability of Brain-to-Speech Models Across Speech Modes
di: Maghsoudi, Maryam, et al.
Pubblicazione: (2026)
di: Maghsoudi, Maryam, et al.
Pubblicazione: (2026)
Aligning Attention with Human Rationales for Self-Explaining Hate Speech Detection
di: Eilertsen, Brage, et al.
Pubblicazione: (2025)
di: Eilertsen, Brage, et al.
Pubblicazione: (2025)
Fine-Grained Interpretation of Political Opinions in Large Language Models
di: Hu, Jingyu, et al.
Pubblicazione: (2025)
di: Hu, Jingyu, et al.
Pubblicazione: (2025)
AutoDetect: Towards a Unified Framework for Automated Weakness Detection in Large Language Models
di: Cheng, Jiale, et al.
Pubblicazione: (2024)
di: Cheng, Jiale, et al.
Pubblicazione: (2024)
Towards Inference-time Category-wise Safety Steering for Large Language Models
di: Bhattacharjee, Amrita, et al.
Pubblicazione: (2024)
di: Bhattacharjee, Amrita, et al.
Pubblicazione: (2024)
Rethinking Interpretability in the Era of Large Language Models
di: Singh, Chandan, et al.
Pubblicazione: (2024)
di: Singh, Chandan, et al.
Pubblicazione: (2024)
OSPC: Artificial VLM Features for Hateful Meme Detection
di: Grönquist, Peter
Pubblicazione: (2024)
di: Grönquist, Peter
Pubblicazione: (2024)
Binary Autoencoder for Mechanistic Interpretability of Large Language Models
di: Cho, Hakaze, et al.
Pubblicazione: (2025)
di: Cho, Hakaze, et al.
Pubblicazione: (2025)
Hate Speech Detection using Large Language Models with Data Augmentation and Feature Enhancement
di: Nge, Brian Jing Hong, et al.
Pubblicazione: (2026)
di: Nge, Brian Jing Hong, et al.
Pubblicazione: (2026)
Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content
di: Stepanov, Ihor, et al.
Pubblicazione: (2026)
di: Stepanov, Ihor, et al.
Pubblicazione: (2026)
HateTinyLLM : Hate Speech Detection Using Tiny Large Language Models
di: Sen, Tanmay, et al.
Pubblicazione: (2024)
di: Sen, Tanmay, et al.
Pubblicazione: (2024)
A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models
di: Shu, Dong, et al.
Pubblicazione: (2025)
di: Shu, Dong, et al.
Pubblicazione: (2025)
SelfIE: Self-Interpretation of Large Language Model Embeddings
di: Chen, Haozhe, et al.
Pubblicazione: (2024)
di: Chen, Haozhe, et al.
Pubblicazione: (2024)
TracrBench: Generating Interpretability Testbeds with Large Language Models
di: Thurnherr, Hannes, et al.
Pubblicazione: (2024)
di: Thurnherr, Hannes, et al.
Pubblicazione: (2024)
HatePrototypes: Interpretable and Transferable Representations for Implicit and Explicit Hate Speech Detection
di: Proskurina, Irina, et al.
Pubblicazione: (2025)
di: Proskurina, Irina, et al.
Pubblicazione: (2025)
Causal Feature Selection for Responsible Machine Learning
di: Moraffah, Raha, et al.
Pubblicazione: (2024)
di: Moraffah, Raha, et al.
Pubblicazione: (2024)
Do LLMs Understand Ambiguity in Text? A Case Study in Open-world Question Answering
di: Keluskar, Aryan, et al.
Pubblicazione: (2024)
di: Keluskar, Aryan, et al.
Pubblicazione: (2024)
Extracting Heuristics from Large Language Models for Reward Shaping in Reinforcement Learning
di: Bhambri, Siddhant, et al.
Pubblicazione: (2024)
di: Bhambri, Siddhant, et al.
Pubblicazione: (2024)
Robust Adaptation of Large Multimodal Models for Retrieval Augmented Hateful Meme Detection
di: Mei, Jingbiao, et al.
Pubblicazione: (2025)
di: Mei, Jingbiao, et al.
Pubblicazione: (2025)
Generating Multiple-Choice Knowledge Questions with Interpretable Difficulty Estimation using Knowledge Graphs and Large Language Models
di: Şakiroğlu, Mehmet Can, et al.
Pubblicazione: (2026)
di: Şakiroğlu, Mehmet Can, et al.
Pubblicazione: (2026)
Safe-SAIL: Towards a Fine-grained Safety Landscape of Large Language Models via Sparse Autoencoder Interpretation Framework
di: Weng, Jiaqi, et al.
Pubblicazione: (2025)
di: Weng, Jiaqi, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Causality Guided Representation Learning for Cross-Style Hate Speech Detection
di: Zhao, Chengshuai, et al.
Pubblicazione: (2025) -
Harnessing Artificial Intelligence to Combat Online Hate: Exploring the Challenges and Opportunities of Large Language Models in Hate Speech Detection
di: Kumarage, Tharindu, et al.
Pubblicazione: (2024) -
Cross-Platform Hate Speech Detection with Weakly Supervised Causal Disentanglement
di: Sheth, Paras, et al.
Pubblicazione: (2024) -
Adversarial Text Purification: A Large Language Model Approach for Defense
di: Moraffah, Raha, et al.
Pubblicazione: (2024) -
Towards LLM-guided Causal Explainability for Black-box Text Classifiers
di: Bhattacharjee, Amrita, et al.
Pubblicazione: (2023)