SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
Fuente:
arXiv
Salvato in:
| Autori principali: | Pandey, Punya Syon, Le, Hai Son, Bhardwaj, Devansh, Mihalcea, Rada, Jin, Zhijing |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
BinaryPPO: Efficient Policy Optimization for Binary Classification
di: Pandey, Punya Syon, et al.
Pubblicazione: (2026)
di: Pandey, Punya Syon, et al.
Pubblicazione: (2026)
Accidental Vulnerability: Factors in Fine-Tuning that Shift Model Safeguards
di: Pandey, Punya Syon, et al.
Pubblicazione: (2025)
di: Pandey, Punya Syon, et al.
Pubblicazione: (2025)
CORE: Measuring Multi-Agent LLM Interaction Quality under Game-Theoretic Pressures
di: Pandey, Punya Syon, et al.
Pubblicazione: (2025)
di: Pandey, Punya Syon, et al.
Pubblicazione: (2025)
Quriosity: Analyzing Human Questioning Behavior and Causal Inquiry through Curiosity-Driven Queries
di: Ceraolo, Roberto, et al.
Pubblicazione: (2024)
di: Ceraolo, Roberto, et al.
Pubblicazione: (2024)
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
di: Andriushchenko, Maksym, et al.
Pubblicazione: (2024)
di: Andriushchenko, Maksym, et al.
Pubblicazione: (2024)
Voices of Her: Analyzing Gender Differences in the AI Publication World
di: Ding, Yiwen, et al.
Pubblicazione: (2023)
di: Ding, Yiwen, et al.
Pubblicazione: (2023)
ChineseHarm-Bench: A Chinese Harmful Content Detection Benchmark
di: Liu, Kangwei, et al.
Pubblicazione: (2025)
di: Liu, Kangwei, et al.
Pubblicazione: (2025)
Can Large Language Models Infer Causation from Correlation?
di: Jin, Zhijing, et al.
Pubblicazione: (2023)
di: Jin, Zhijing, et al.
Pubblicazione: (2023)
"They are uncultured": Unveiling Covert Harms and Social Threats in LLM Generated Conversations
di: Dammu, Preetam Prabhu Srikar, et al.
Pubblicazione: (2024)
di: Dammu, Preetam Prabhu Srikar, et al.
Pubblicazione: (2024)
When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas
di: Backmann, Steffen, et al.
Pubblicazione: (2025)
di: Backmann, Steffen, et al.
Pubblicazione: (2025)
Evaluating Cooperation in LLM Social Groups through Elected Leadership
di: Faulkner, Ryan, et al.
Pubblicazione: (2026)
di: Faulkner, Ryan, et al.
Pubblicazione: (2026)
IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures
di: Gringras, David
Pubblicazione: (2026)
di: Gringras, David
Pubblicazione: (2026)
Cross-cultural Inspiration Detection and Analysis in Real and LLM-generated Social Media Data
di: Ignat, Oana, et al.
Pubblicazione: (2024)
di: Ignat, Oana, et al.
Pubblicazione: (2024)
A Generative Approach to LLM Harmfulness Mitigation with Red Flag Tokens
di: Dobre, David, et al.
Pubblicazione: (2025)
di: Dobre, David, et al.
Pubblicazione: (2025)
Unintended Harms of Value-Aligned LLMs: Psychological and Empirical Insights
di: Choi, Sooyung, et al.
Pubblicazione: (2025)
di: Choi, Sooyung, et al.
Pubblicazione: (2025)
Implicit Personalization in Language Models: A Systematic Study
di: Jin, Zhijing, et al.
Pubblicazione: (2024)
di: Jin, Zhijing, et al.
Pubblicazione: (2024)
Harmful Intent as a Geometrically Recoverable Feature of LLM Residual Streams
di: Llorente-Saguer, Isaac
Pubblicazione: (2026)
di: Llorente-Saguer, Isaac
Pubblicazione: (2026)
Toxicity Detection Should Measure Contextual Harm, Not Text-Intrinsic Badness
di: Berezin, Sergei, et al.
Pubblicazione: (2025)
di: Berezin, Sergei, et al.
Pubblicazione: (2025)
The Multilingual Alignment Prism: Aligning Global and Local Preferences to Reduce Harm
di: Aakanksha, et al.
Pubblicazione: (2024)
di: Aakanksha, et al.
Pubblicazione: (2024)
Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
di: Sheshadri, Abhay, et al.
Pubblicazione: (2024)
di: Sheshadri, Abhay, et al.
Pubblicazione: (2024)
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
di: Mazeika, Mantas, et al.
Pubblicazione: (2024)
di: Mazeika, Mantas, et al.
Pubblicazione: (2024)
Deception Detection from Linguistic and Physiological Data Streams Using Bimodal Convolutional Neural Networks
di: Li, Panfeng, et al.
Pubblicazione: (2023)
di: Li, Panfeng, et al.
Pubblicazione: (2023)
SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
di: Chaudhary, Maheep, et al.
Pubblicazione: (2025)
di: Chaudhary, Maheep, et al.
Pubblicazione: (2025)
Towards Generalizable Generic Harmful Speech Datasets for Implicit Hate Speech Detection
di: Almohaimeed, Saad, et al.
Pubblicazione: (2025)
di: Almohaimeed, Saad, et al.
Pubblicazione: (2025)
Causality for Natural Language Processing
di: Jin, Zhijing
Pubblicazione: (2025)
di: Jin, Zhijing
Pubblicazione: (2025)
Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content
di: Stepanov, Ihor, et al.
Pubblicazione: (2026)
di: Stepanov, Ihor, et al.
Pubblicazione: (2026)
Laissez-Faire Harms: Algorithmic Biases in Generative Language Models
di: Shieh, Evan, et al.
Pubblicazione: (2024)
di: Shieh, Evan, et al.
Pubblicazione: (2024)
Beyond Behaviorist Representational Harms: A Plan for Measurement and Mitigation
di: Chien, Jennifer, et al.
Pubblicazione: (2024)
di: Chien, Jennifer, et al.
Pubblicazione: (2024)
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation
di: Kim, Eunsu, et al.
Pubblicazione: (2025)
di: Kim, Eunsu, et al.
Pubblicazione: (2025)
Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple Interactions
di: Chan, Yik Siu, et al.
Pubblicazione: (2025)
di: Chan, Yik Siu, et al.
Pubblicazione: (2025)
Causally Testing Gender Bias in LLMs: A Case Study on Occupational Bias
di: Chen, Yuen, et al.
Pubblicazione: (2022)
di: Chen, Yuen, et al.
Pubblicazione: (2022)
Objective Matters: Fine-Tuning Objectives Shape Safety, Robustness, and Persona Drift
di: Vennemeyer, Daniel, et al.
Pubblicazione: (2026)
di: Vennemeyer, Daniel, et al.
Pubblicazione: (2026)
The Geometry of Harmful Intent: Training-Free Anomaly Detection via Angular Deviation in LLM Residual Streams
di: Llorente-Saguer, Isaac
Pubblicazione: (2026)
di: Llorente-Saguer, Isaac
Pubblicazione: (2026)
Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap
di: Qi, Xuan, et al.
Pubblicazione: (2025)
di: Qi, Xuan, et al.
Pubblicazione: (2025)
MuPHI: Learning Implicit Multimodal Harm Reasoning via Semantically Grounded Reward Optimization
di: Saha, Anisha, et al.
Pubblicazione: (2026)
di: Saha, Anisha, et al.
Pubblicazione: (2026)
HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
di: Yang, Langqi, et al.
Pubblicazione: (2025)
di: Yang, Langqi, et al.
Pubblicazione: (2025)
Self-HarmLLM: Can Large Language Model Harm Itself?
di: Kim, Heehwan, et al.
Pubblicazione: (2025)
di: Kim, Heehwan, et al.
Pubblicazione: (2025)
Eliciting Harmful Capabilities by Fine-Tuning On Safeguarded Outputs
di: Kaunismaa, Jackson, et al.
Pubblicazione: (2026)
di: Kaunismaa, Jackson, et al.
Pubblicazione: (2026)
Virus: Harmful Fine-tuning Attack for Large Language Models Bypassing Guardrail Moderation
di: Huang, Tiansheng, et al.
Pubblicazione: (2025)
di: Huang, Tiansheng, et al.
Pubblicazione: (2025)
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency
di: Li, Ruixiao, et al.
Pubblicazione: (2025)
di: Li, Ruixiao, et al.
Pubblicazione: (2025)
Documenti analoghi
-
BinaryPPO: Efficient Policy Optimization for Binary Classification
di: Pandey, Punya Syon, et al.
Pubblicazione: (2026) -
Accidental Vulnerability: Factors in Fine-Tuning that Shift Model Safeguards
di: Pandey, Punya Syon, et al.
Pubblicazione: (2025) -
CORE: Measuring Multi-Agent LLM Interaction Quality under Game-Theoretic Pressures
di: Pandey, Punya Syon, et al.
Pubblicazione: (2025) -
Quriosity: Analyzing Human Questioning Behavior and Causal Inquiry through Curiosity-Driven Queries
di: Ceraolo, Roberto, et al.
Pubblicazione: (2024) -
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
di: Andriushchenko, Maksym, et al.
Pubblicazione: (2024)