Moderating Harm: Benchmarking Large Language Models for Cyberbullying Detection in YouTube Comments
Fuente:
arXiv
Saved in:
| Main Author: | Muminovic, Amel |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Large Language Models for Toxic Language Detection in Low-Resource Balkan Languages
by: Muminovic, Amel, et al.
Published: (2025)
by: Muminovic, Amel, et al.
Published: (2025)
YouTube Comments Decoded: Leveraging LLMs for Low Resource Language Classification
by: Deroy, Aniket, et al.
Published: (2024)
by: Deroy, Aniket, et al.
Published: (2024)
DIVERSE: A Dataset of YouTube Video Comment Stances with a Data Programming Model
by: Cruickshank, Iain J., et al.
Published: (2024)
by: Cruickshank, Iain J., et al.
Published: (2024)
DariMis: Harm-Aware Modeling for Dari Misinformation Detection on YouTube
by: Baktash, Jawid Ahmad, et al.
Published: (2026)
by: Baktash, Jawid Ahmad, et al.
Published: (2026)
A Novel BERT-based Classifier to Detect Political Leaning of YouTube Videos based on their Titles
by: AlDahoul, Nouar, et al.
Published: (2024)
by: AlDahoul, Nouar, et al.
Published: (2024)
Mapping Controversies Using Artificial Intelligence: An Analysis of the Hamas-Israel Conflict on YouTube
by: Lopez, Victor Manuel Hernandez, et al.
Published: (2025)
by: Lopez, Victor Manuel Hernandez, et al.
Published: (2025)
Harmful YouTube Video Detection: A Taxonomy of Online Harm and MLLMs as Alternative Annotators
by: Jo, Claire Wonjeong, et al.
Published: (2024)
by: Jo, Claire Wonjeong, et al.
Published: (2024)
Truth Sleuth and Trend Bender: AI Agents to fact-check YouTube videos and influence opinions
by: Logé, Cécile, et al.
Published: (2025)
by: Logé, Cécile, et al.
Published: (2025)
Self-HarmLLM: Can Large Language Model Harm Itself?
by: Kim, Heehwan, et al.
Published: (2025)
by: Kim, Heehwan, et al.
Published: (2025)
The Use of a Large Language Model for Cyberbullying Detection
by: Ogunleye, Bayode, et al.
Published: (2024)
by: Ogunleye, Bayode, et al.
Published: (2024)
Multimodal Analysis of State-Funded News Coverage of the Israel-Hamas War on YouTube Shorts
by: Miehling, Daniel, et al.
Published: (2026)
by: Miehling, Daniel, et al.
Published: (2026)
A "Perspectival" Mirror of the Elephant: Investigating Language Bias on Google, ChatGPT, YouTube, and Wikipedia
by: Luo, Queenie, et al.
Published: (2023)
by: Luo, Queenie, et al.
Published: (2023)
GODBench: A Benchmark for Multimodal Large Language Models in Video Comment Art
by: Lei, Yiming, et al.
Published: (2025)
by: Lei, Yiming, et al.
Published: (2025)
Booster: Tackling Harmful Fine-tuning for Large Language Models via Attenuating Harmful Perturbation
by: Huang, Tiansheng, et al.
Published: (2024)
by: Huang, Tiansheng, et al.
Published: (2024)
Virus: Harmful Fine-tuning Attack for Large Language Models Bypassing Guardrail Moderation
by: Huang, Tiansheng, et al.
Published: (2025)
by: Huang, Tiansheng, et al.
Published: (2025)
Towards Explainable Harmful Meme Detection through Multimodal Debate between Large Language Models
by: Lin, Hongzhan, et al.
Published: (2024)
by: Lin, Hongzhan, et al.
Published: (2024)
Chinese Cyberbullying Detection: Dataset, Method, and Validation
by: Zhu, Yi, et al.
Published: (2025)
by: Zhu, Yi, et al.
Published: (2025)
Moral Outrage Shapes Commitments Beyond Attention: Multimodal Moral Emotions on YouTube in Korea and the US
by: Park, Seongchan, et al.
Published: (2026)
by: Park, Seongchan, et al.
Published: (2026)
Advancing Harmful Content Detection in Organizational Research: Integrating Large Language Models with Elo Rating System
by: Akben, Mustafa, et al.
Published: (2025)
by: Akben, Mustafa, et al.
Published: (2025)
OSPC: Detecting Harmful Memes with Large Language Model as a Catalyst
by: Cao, Jingtao, et al.
Published: (2024)
by: Cao, Jingtao, et al.
Published: (2024)
Identifying False Content and Hate Speech in Sinhala YouTube Videos by Analyzing the Audio
by: Wickramaarachchi, W. A. K. M., et al.
Published: (2024)
by: Wickramaarachchi, W. A. K. M., et al.
Published: (2024)
Who Shapes Brazil's Vaccine Debate? Semi-Supervised Modeling of Stance and Polarization in YouTube's Media Ecosystem
by: de Oliveira, Geovana S., et al.
Published: (2026)
by: de Oliveira, Geovana S., et al.
Published: (2026)
You've Changed: Detecting Modification of Black-Box Large Language Models
by: Dima, Alden, et al.
Published: (2025)
by: Dima, Alden, et al.
Published: (2025)
HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
by: Yang, Langqi, et al.
Published: (2025)
by: Yang, Langqi, et al.
Published: (2025)
AD-LLM: Benchmarking Large Language Models for Anomaly Detection
by: Yang, Tiankai, et al.
Published: (2024)
by: Yang, Tiankai, et al.
Published: (2024)
Large Language Models are Vulnerable to Bait-and-Switch Attacks for Generating Harmful Content
by: Bianchi, Federico, et al.
Published: (2024)
by: Bianchi, Federico, et al.
Published: (2024)
Harm or Humor: A Multimodal, Multilingual Benchmark for Overt and Covert Harmful Humor
by: Sharshar, Ahmed, et al.
Published: (2026)
by: Sharshar, Ahmed, et al.
Published: (2026)
AdamMeme: Adaptively Probe the Reasoning Capacity of Multimodal Large Language Models on Harmfulness
by: Chen, Zixin, et al.
Published: (2025)
by: Chen, Zixin, et al.
Published: (2025)
VaccineRAG: Boosting Multimodal Large Language Models' Immunity to Harmful RAG Samples
by: Sun, Qixin, et al.
Published: (2025)
by: Sun, Qixin, et al.
Published: (2025)
Context-Aware Content Moderation for German Newspaper Comments
by: Krejca, Felix, et al.
Published: (2025)
by: Krejca, Felix, et al.
Published: (2025)
PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
by: Li, Jing-Jing, et al.
Published: (2026)
by: Li, Jing-Jing, et al.
Published: (2026)
YT-30M: A multi-lingual multi-category dataset of YouTube comments
by: Dutta, Hridoy Sankar
Published: (2024)
by: Dutta, Hridoy Sankar
Published: (2024)
Towards Safer Social Media Platforms: Scalable and Performant Few-Shot Harmful Content Moderation Using Large Language Models
by: Bonagiri, Akash, et al.
Published: (2025)
by: Bonagiri, Akash, et al.
Published: (2025)
FairBelief -- Assessing Harmful Beliefs in Language Models
by: Setzu, Mattia, et al.
Published: (2024)
by: Setzu, Mattia, et al.
Published: (2024)
Deep Learning Based Cyberbullying Detection in Bangla Language
by: Nath, Sristy Shidul, et al.
Published: (2024)
by: Nath, Sristy Shidul, et al.
Published: (2024)
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
by: Andriushchenko, Maksym, et al.
Published: (2024)
by: Andriushchenko, Maksym, et al.
Published: (2024)
MemeArena: Automating Context-Aware Unbiased Evaluation of Harmfulness Understanding for Multimodal Large Language Models
by: Chen, Zixin, et al.
Published: (2025)
by: Chen, Zixin, et al.
Published: (2025)
Panacea: Mitigating Harmful Fine-tuning for Large Language Models via Post-fine-tuning Perturbation
by: Wang, Yibo, et al.
Published: (2025)
by: Wang, Yibo, et al.
Published: (2025)
Nevermind: Instruction Override and Moderation in Large Language Models
by: Kim, Edward
Published: (2024)
by: Kim, Edward
Published: (2024)
Red Lines and Grey Zones in the Fog of War: Benchmarking Legal Risk, Moral Harm, and Regional Bias in Large Language Model Military Decision-Making
by: Drinkall, Toby
Published: (2025)
by: Drinkall, Toby
Published: (2025)
Similar Items
-
Large Language Models for Toxic Language Detection in Low-Resource Balkan Languages
by: Muminovic, Amel, et al.
Published: (2025) -
YouTube Comments Decoded: Leveraging LLMs for Low Resource Language Classification
by: Deroy, Aniket, et al.
Published: (2024) -
DIVERSE: A Dataset of YouTube Video Comment Stances with a Data Programming Model
by: Cruickshank, Iain J., et al.
Published: (2024) -
DariMis: Harm-Aware Modeling for Dari Misinformation Detection on YouTube
by: Baktash, Jawid Ahmad, et al.
Published: (2026) -
A Novel BERT-based Classifier to Detect Political Leaning of YouTube Videos based on their Titles
by: AlDahoul, Nouar, et al.
Published: (2024)