HarmPot: An Annotation Framework for Evaluating Offline Harm Potential of Social Media Text
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kumar, Ritesh, Bhalla, Ojaswee, Vanthi, Madhu, Wani, Shehlat Maknoon, Singh, Siddharth |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
NJUST-KMG at TRAC-2024 Tasks 1 and 2: Offline Harm Potential Identification
von: Wang, Jingyuan, et al.
Veröffentlicht: (2024)
von: Wang, Jingyuan, et al.
Veröffentlicht: (2024)
SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
von: Pandey, Punya Syon, et al.
Veröffentlicht: (2025)
von: Pandey, Punya Syon, et al.
Veröffentlicht: (2025)
Careless Whisper: Speech-to-Text Hallucination Harms
von: Koenecke, Allison, et al.
Veröffentlicht: (2024)
von: Koenecke, Allison, et al.
Veröffentlicht: (2024)
HarmLevelBench: Evaluating Harm-Level Compliance and the Impact of Quantization on Model Alignment
von: Belkhiter, Yannis, et al.
Veröffentlicht: (2024)
von: Belkhiter, Yannis, et al.
Veröffentlicht: (2024)
Improving Harmful Text Detection with Joint Retrieval and External Knowledge
von: Yu, Zidong, et al.
Veröffentlicht: (2025)
von: Yu, Zidong, et al.
Veröffentlicht: (2025)
StopHC: A Harmful Content Detection and Mitigation Architecture for Social Media Platforms
von: Truică, Ciprian-Octavian, et al.
Veröffentlicht: (2024)
von: Truică, Ciprian-Octavian, et al.
Veröffentlicht: (2024)
PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
von: Li, Jing-Jing, et al.
Veröffentlicht: (2026)
von: Li, Jing-Jing, et al.
Veröffentlicht: (2026)
An Evaluation of LLMs for Detecting Harmful Computing Terms
von: Jacas, Joshua, et al.
Veröffentlicht: (2025)
von: Jacas, Joshua, et al.
Veröffentlicht: (2025)
Directed Social Regard: Surfacing Targeted Advocacy, Opposition, Aid, Harms, and Victimization in Online Media
von: Friedman, Scott, et al.
Veröffentlicht: (2026)
von: Friedman, Scott, et al.
Veröffentlicht: (2026)
HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
von: Yang, Langqi, et al.
Veröffentlicht: (2025)
von: Yang, Langqi, et al.
Veröffentlicht: (2025)
Self-HarmLLM: Can Large Language Model Harm Itself?
von: Kim, Heehwan, et al.
Veröffentlicht: (2025)
von: Kim, Heehwan, et al.
Veröffentlicht: (2025)
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
Harm or Humor: A Multimodal, Multilingual Benchmark for Overt and Covert Harmful Humor
von: Sharshar, Ahmed, et al.
Veröffentlicht: (2026)
von: Sharshar, Ahmed, et al.
Veröffentlicht: (2026)
"They are uncultured": Unveiling Covert Harms and Social Threats in LLM Generated Conversations
von: Dammu, Preetam Prabhu Srikar, et al.
Veröffentlicht: (2024)
von: Dammu, Preetam Prabhu Srikar, et al.
Veröffentlicht: (2024)
The Hidden Language of Harm: Examining the Role of Emojis in Harmful Online Communication and Content Moderation
von: Zhou, Yuhang, et al.
Veröffentlicht: (2025)
von: Zhou, Yuhang, et al.
Veröffentlicht: (2025)
HarmTransform: Transforming Explicit Harmful Queries into Stealthy via Multi-Agent Debate
von: Zhu, Shenzhe
Veröffentlicht: (2025)
von: Zhu, Shenzhe
Veröffentlicht: (2025)
Can Editing LLMs Inject Harm?
von: Chen, Canyu, et al.
Veröffentlicht: (2024)
von: Chen, Canyu, et al.
Veröffentlicht: (2024)
The Psychosocial Impacts of Generative AI Harms
von: Vassel, Faye-Marie, et al.
Veröffentlicht: (2024)
von: Vassel, Faye-Marie, et al.
Veröffentlicht: (2024)
LLMs Encode Harmfulness and Refusal Separately
von: Zhao, Jiachen, et al.
Veröffentlicht: (2025)
von: Zhao, Jiachen, et al.
Veröffentlicht: (2025)
Lost in Translation? A Comparative Study on the Cross-Lingual Transfer of Composite Harms
von: Shukla, Vaibhav, et al.
Veröffentlicht: (2026)
von: Shukla, Vaibhav, et al.
Veröffentlicht: (2026)
Between Help and Harm: An Evaluation of Mental Health Crisis Handling by LLMs
von: Arnaiz-Rodriguez, Adrian, et al.
Veröffentlicht: (2025)
von: Arnaiz-Rodriguez, Adrian, et al.
Veröffentlicht: (2025)
Booster: Tackling Harmful Fine-tuning for Large Language Models via Attenuating Harmful Perturbation
von: Huang, Tiansheng, et al.
Veröffentlicht: (2024)
von: Huang, Tiansheng, et al.
Veröffentlicht: (2024)
HALF: Harm-Aware LLM Fairness Evaluation Aligned with Deployment
von: Mekky, Ali, et al.
Veröffentlicht: (2025)
von: Mekky, Ali, et al.
Veröffentlicht: (2025)
Re-ranking Using Large Language Models for Mitigating Exposure to Harmful Content on Social Media Platforms
von: Oak, Rajvardhan, et al.
Veröffentlicht: (2025)
von: Oak, Rajvardhan, et al.
Veröffentlicht: (2025)
From Representational Harms to Quality-of-Service Harms: A Case Study on Llama 2 Safety Safeguards
von: Chehbouni, Khaoula, et al.
Veröffentlicht: (2024)
von: Chehbouni, Khaoula, et al.
Veröffentlicht: (2024)
Target Span Detection for Implicit Harmful Content
von: Jafari, Nazanin, et al.
Veröffentlicht: (2024)
von: Jafari, Nazanin, et al.
Veröffentlicht: (2024)
`For Argument's Sake, Show Me How to Harm Myself!': Jailbreaking LLMs in Suicide and Self-Harm Contexts
von: Schoene, Annika M, et al.
Veröffentlicht: (2025)
von: Schoene, Annika M, et al.
Veröffentlicht: (2025)
Toxicity Detection Should Measure Contextual Harm, Not Text-Intrinsic Badness
von: Berezin, Sergei, et al.
Veröffentlicht: (2025)
von: Berezin, Sergei, et al.
Veröffentlicht: (2025)
Expected Harm: Rethinking Safety Evaluation of (Mis)Aligned LLMs
von: Chen, Yen-Shan, et al.
Veröffentlicht: (2026)
von: Chen, Yen-Shan, et al.
Veröffentlicht: (2026)
Deep Research Brings Deeper Harm
von: Chen, Shuo, et al.
Veröffentlicht: (2025)
von: Chen, Shuo, et al.
Veröffentlicht: (2025)
ToxiCraft: A Novel Framework for Synthetic Generation of Harmful Information
von: Hui, Zheng, et al.
Veröffentlicht: (2024)
von: Hui, Zheng, et al.
Veröffentlicht: (2024)
REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM
von: Jindal, Madhur, et al.
Veröffentlicht: (2025)
von: Jindal, Madhur, et al.
Veröffentlicht: (2025)
Harmful Suicide Content Detection
von: Park, Kyumin, et al.
Veröffentlicht: (2024)
von: Park, Kyumin, et al.
Veröffentlicht: (2024)
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
von: Mazeika, Mantas, et al.
Veröffentlicht: (2024)
von: Mazeika, Mantas, et al.
Veröffentlicht: (2024)
Red Teaming Multimodal Language Models: Evaluating Harm Across Prompt Modalities and Models
von: Van Doren, Madison, et al.
Veröffentlicht: (2025)
von: Van Doren, Madison, et al.
Veröffentlicht: (2025)
ChineseHarm-Bench: A Chinese Harmful Content Detection Benchmark
von: Liu, Kangwei, et al.
Veröffentlicht: (2025)
von: Liu, Kangwei, et al.
Veröffentlicht: (2025)
Taxonomizing Representational Harms using Speech Act Theory
von: Corvi, Emily, et al.
Veröffentlicht: (2025)
von: Corvi, Emily, et al.
Veröffentlicht: (2025)
LLM-based Semantic Augmentation for Harmful Content Detection
von: Meguellati, Elyas, et al.
Veröffentlicht: (2025)
von: Meguellati, Elyas, et al.
Veröffentlicht: (2025)
MIND: A Multi-agent Framework for Zero-shot Harmful Meme Detection
von: Liu, Ziyan, et al.
Veröffentlicht: (2025)
von: Liu, Ziyan, et al.
Veröffentlicht: (2025)
Towards Comprehensive Detection of Chinese Harmful Memes
von: Lu, Junyu, et al.
Veröffentlicht: (2024)
von: Lu, Junyu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
NJUST-KMG at TRAC-2024 Tasks 1 and 2: Offline Harm Potential Identification
von: Wang, Jingyuan, et al.
Veröffentlicht: (2024) -
SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
von: Pandey, Punya Syon, et al.
Veröffentlicht: (2025) -
Careless Whisper: Speech-to-Text Hallucination Harms
von: Koenecke, Allison, et al.
Veröffentlicht: (2024) -
HarmLevelBench: Evaluating Harm-Level Compliance and the Impact of Quantization on Model Alignment
von: Belkhiter, Yannis, et al.
Veröffentlicht: (2024) -
Improving Harmful Text Detection with Joint Retrieval and External Knowledge
von: Yu, Zidong, et al.
Veröffentlicht: (2025)