Cross-Platform Hate Speech Detection with Weakly Supervised Causal Disentanglement
Fuente:
arXiv
Saved in:
| Main Authors: | Sheth, Paras, Kumarage, Tharindu, Moraffah, Raha, Chadha, Aman, Liu, Huan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Survey of AI-generated Text Forensic Systems: Detection, Attribution, and Characterization
by: Kumarage, Tharindu, et al.
Published: (2024)
by: Kumarage, Tharindu, et al.
Published: (2024)
Causality Guided Representation Learning for Cross-Style Hate Speech Detection
by: Zhao, Chengshuai, et al.
Published: (2025)
by: Zhao, Chengshuai, et al.
Published: (2025)
Can Large Language Models Infer Causal Relationships from Real-World Text?
by: Saklad, Ryan, et al.
Published: (2025)
by: Saklad, Ryan, et al.
Published: (2025)
Causal Feature Selection for Responsible Machine Learning
by: Moraffah, Raha, et al.
Published: (2024)
by: Moraffah, Raha, et al.
Published: (2024)
Towards Interpretable Hate Speech Detection using Large Language Model-extracted Rationales
by: Nirmal, Ayushi, et al.
Published: (2024)
by: Nirmal, Ayushi, et al.
Published: (2024)
Exploiting Class Probabilities for Black-box Sentence-level Attacks
by: Moraffah, Raha, et al.
Published: (2024)
by: Moraffah, Raha, et al.
Published: (2024)
Towards LLM-guided Causal Explainability for Black-box Text Classifiers
by: Bhattacharjee, Amrita, et al.
Published: (2023)
by: Bhattacharjee, Amrita, et al.
Published: (2023)
EAGLE: A Domain Generalization Framework for AI-generated Text Detection
by: Bhattacharjee, Amrita, et al.
Published: (2024)
by: Bhattacharjee, Amrita, et al.
Published: (2024)
Can Knowledge Graphs Reduce Hallucinations in LLMs? : A Survey
by: Agrawal, Garima, et al.
Published: (2023)
by: Agrawal, Garima, et al.
Published: (2023)
Harnessing Artificial Intelligence to Combat Online Hate: Exploring the Challenges and Opportunities of Large Language Models in Hate Speech Detection
by: Kumarage, Tharindu, et al.
Published: (2024)
by: Kumarage, Tharindu, et al.
Published: (2024)
Zero-shot LLM-guided Counterfactual Generation: A Case Study on NLP Model Evaluation
by: Bhattacharjee, Amrita, et al.
Published: (2024)
by: Bhattacharjee, Amrita, et al.
Published: (2024)
Adversarial Text Purification: A Large Language Model Approach for Defense
by: Moraffah, Raha, et al.
Published: (2024)
by: Moraffah, Raha, et al.
Published: (2024)
Investigating Annotator Bias in Large Language Models for Hate Speech Detection
by: Das, Amit, et al.
Published: (2024)
by: Das, Amit, et al.
Published: (2024)
A Generative Approach to Surrogate-based Black-box Attacks
by: Moraffah, Raha, et al.
Published: (2024)
by: Moraffah, Raha, et al.
Published: (2024)
Advancing Hate Speech Detection with Transformers: Insights from the MetaHate
by: Chapagain, Santosh, et al.
Published: (2025)
by: Chapagain, Santosh, et al.
Published: (2025)
HateTinyLLM : Hate Speech Detection Using Tiny Large Language Models
by: Sen, Tanmay, et al.
Published: (2024)
by: Sen, Tanmay, et al.
Published: (2024)
Disentangling Latent Shifts of In-Context Learning with Weak Supervision
by: Jukić, Josip, et al.
Published: (2024)
by: Jukić, Josip, et al.
Published: (2024)
Hate Speech Detection with Generalizable Target-aware Fairness
by: Chen, Tong, et al.
Published: (2024)
by: Chen, Tong, et al.
Published: (2024)
Hate Speech Detection and Classification in Amharic Text with Deep Learning
by: Gashe, Samuel Minale, et al.
Published: (2024)
by: Gashe, Samuel Minale, et al.
Published: (2024)
An Effective, Robust and Fairness-aware Hate Speech Detection Framework
by: Mou, Guanyi, et al.
Published: (2024)
by: Mou, Guanyi, et al.
Published: (2024)
Aligning Attention with Human Rationales for Self-Explaining Hate Speech Detection
by: Eilertsen, Brage, et al.
Published: (2025)
by: Eilertsen, Brage, et al.
Published: (2025)
A Survey on Automatic Online Hate Speech Detection in Low-Resource Languages
by: Das, Susmita, et al.
Published: (2024)
by: Das, Susmita, et al.
Published: (2024)
Guided Distant Supervision for Multilingual Relation Extraction Data: Adapting to a New Language
by: Plum, Alistair, et al.
Published: (2024)
by: Plum, Alistair, et al.
Published: (2024)
BOISHOMMO: Holistic Approach for Bangla Hate Speech
by: Kafi, Md Abdullah Al, et al.
Published: (2025)
by: Kafi, Md Abdullah Al, et al.
Published: (2025)
Towards Generalizable Generic Harmful Speech Datasets for Implicit Hate Speech Detection
by: Almohaimeed, Saad, et al.
Published: (2025)
by: Almohaimeed, Saad, et al.
Published: (2025)
Deep Learning Approaches for Detecting Adversarial Cyberbullying and Hate Speech in Social Networks
by: Azumah, Sylvia Worlali, et al.
Published: (2024)
by: Azumah, Sylvia Worlali, et al.
Published: (2024)
I Can't Believe It's Not Robust: Catastrophic Collapse of Safety Classifiers under Embedding Drift
by: Sahoo, Subramanyam, et al.
Published: (2026)
by: Sahoo, Subramanyam, et al.
Published: (2026)
On Importance of Code-Mixed Embeddings for Hate Speech Identification
by: Jagdale, Shruti, et al.
Published: (2024)
by: Jagdale, Shruti, et al.
Published: (2024)
Are Small Language Models Ready to Compete with Large Language Models for Practical Applications?
by: Sinha, Neelabh, et al.
Published: (2024)
by: Sinha, Neelabh, et al.
Published: (2024)
SWE2: SubWord Enriched and Significant Word Emphasized Framework for Hate Speech Detection
by: Mou, Guanyi, et al.
Published: (2024)
by: Mou, Guanyi, et al.
Published: (2024)
Vicarious Offense and Noise Audit of Offensive Speech Classifiers: Unifying Human and Machine Disagreement on What is Offensive
by: Weerasooriya, Tharindu Cyril, et al.
Published: (2023)
by: Weerasooriya, Tharindu Cyril, et al.
Published: (2023)
Transformers and Ensemble methods: A solution for Hate Speech Detection in Arabic languages
by: de Paula, Angel Felipe Magnossão, et al.
Published: (2023)
by: de Paula, Angel Felipe Magnossão, et al.
Published: (2023)
Improving Hate Speech Classification with Cross-Taxonomy Dataset Integration
by: Fillies, Jan, et al.
Published: (2025)
by: Fillies, Jan, et al.
Published: (2025)
LLMsAgainstHate @ NLU of Devanagari Script Languages 2025: Hate Speech Detection and Target Identification in Devanagari Languages via Parameter Efficient Fine-Tuning of LLMs
by: Sidibomma, Rushendra, et al.
Published: (2024)
by: Sidibomma, Rushendra, et al.
Published: (2024)
Overview of Factify5WQA: Fact Verification through 5W Question-Answering
by: Suresh, Suryavardan, et al.
Published: (2024)
by: Suresh, Suryavardan, et al.
Published: (2024)
When Shallow Wins: Silent Failures and the Depth-Accuracy Paradox in Latent Reasoning
by: Sahoo, Subramanyam, et al.
Published: (2026)
by: Sahoo, Subramanyam, et al.
Published: (2026)
SAHOO: Safeguarded Alignment for High-Order Optimization Objectives in Recursive Self-Improvement
by: Sahoo, Subramanyam, et al.
Published: (2026)
by: Sahoo, Subramanyam, et al.
Published: (2026)
Guiding Vision-Language Model Selection for Visual Question-Answering Across Tasks, Domains, and Knowledge Types
by: Sinha, Neelabh, et al.
Published: (2024)
by: Sinha, Neelabh, et al.
Published: (2024)
From BERT to Qwen: Hate Detection across architectures
by: Mon, Ariadna, et al.
Published: (2025)
by: Mon, Ariadna, et al.
Published: (2025)
Transfer Learning via Lexical Relatedness: A Sarcasm and Hate Speech Case Study
by: Cabrera, Angelly, et al.
Published: (2025)
by: Cabrera, Angelly, et al.
Published: (2025)
Similar Items
-
A Survey of AI-generated Text Forensic Systems: Detection, Attribution, and Characterization
by: Kumarage, Tharindu, et al.
Published: (2024) -
Causality Guided Representation Learning for Cross-Style Hate Speech Detection
by: Zhao, Chengshuai, et al.
Published: (2025) -
Can Large Language Models Infer Causal Relationships from Real-World Text?
by: Saklad, Ryan, et al.
Published: (2025) -
Causal Feature Selection for Responsible Machine Learning
by: Moraffah, Raha, et al.
Published: (2024) -
Towards Interpretable Hate Speech Detection using Large Language Model-extracted Rationales
by: Nirmal, Ayushi, et al.
Published: (2024)