Quantifying Harm
Fuente:
arXiv
Saved in:
| Main Authors: | Beckers, Sander, Chockler, Hana, Halpern, Joseph Y. |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Explaining Image Classifiers
by: Chockler, Hana, et al.
Published: (2024)
by: Chockler, Hana, et al.
Published: (2024)
Clustered Policy Decision Ranking
by: Levin, Mark, et al.
Published: (2023)
by: Levin, Mark, et al.
Published: (2023)
Defining and Quantifying Creative Behavior in Popular Image Generators
by: Ramaswamy, Aditi, et al.
Published: (2025)
by: Ramaswamy, Aditi, et al.
Published: (2025)
Nondeterministic Causal Models
by: Beckers, Sander
Published: (2024)
by: Beckers, Sander
Published: (2024)
Large Language Models as Nondeterministic Causal Models
by: Beckers, Sander
Published: (2025)
by: Beckers, Sander
Published: (2025)
Actual Causation and Nondeterministic Causal Models
by: Beckers, Sander
Published: (2025)
by: Beckers, Sander
Published: (2025)
Causal Counterfactuals Reconsidered
by: Beckers, Sander
Published: (2025)
by: Beckers, Sander
Published: (2025)
Sufficient, Necessary and Complete Causal Explanations in Image Classification
by: Kelly, David A, et al.
Published: (2025)
by: Kelly, David A, et al.
Published: (2025)
Mathematical Explanations
by: Halpern, Joseph Y.
Published: (2023)
by: Halpern, Joseph Y.
Published: (2023)
Causal Explanations for Image Classifiers
by: Chockler, Hana, et al.
Published: (2024)
by: Chockler, Hana, et al.
Published: (2024)
Out-of-the-box: Black-box Causal Attacks on Object Detectors
by: Navaratnarajah, Melane, et al.
Published: (2025)
by: Navaratnarajah, Melane, et al.
Published: (2025)
Multiple Different Black Box Explanations for Image Classifiers
by: Chockler, Hana, et al.
Published: (2023)
by: Chockler, Hana, et al.
Published: (2023)
It's a Feature, Not a Bug: Measuring Creative Fluidity in Image Generators
by: Ramaswamy, Aditi, et al.
Published: (2024)
by: Ramaswamy, Aditi, et al.
Published: (2024)
Evaluation of Black-Box XAI Approaches for Predictors of Values of Boolean Formulae
by: Armoni-Friedmann, Stav, et al.
Published: (2025)
by: Armoni-Friedmann, Stav, et al.
Published: (2025)
Counterfactual Influence in Markov Decision Processes
by: Kazemi, Milad, et al.
Published: (2024)
by: Kazemi, Milad, et al.
Published: (2024)
Activation-Deactivation: A General Framework for Robust Post-hoc Explainable AI
by: Chanchal, Akchunya, et al.
Published: (2025)
by: Chanchal, Akchunya, et al.
Published: (2025)
Real-Time Incremental Explanations for Object Detectors in Autonomous Driving
by: Calderón-Peña, Santiago, et al.
Published: (2024)
by: Calderón-Peña, Santiago, et al.
Published: (2024)
Causality Without Causal Models
by: Halpern, Joseph Y., et al.
Published: (2025)
by: Halpern, Joseph Y., et al.
Published: (2025)
Intervention and Conditioning in Causal Bayesian Networks
by: Galhotra, Sainyam, et al.
Published: (2024)
by: Galhotra, Sainyam, et al.
Published: (2024)
Subjective Causality
by: Halpern, Joseph Y., et al.
Published: (2024)
by: Halpern, Joseph Y., et al.
Published: (2024)
SpecReX: Explainable AI for Raman Spectroscopy
by: Blake, Nathan, et al.
Published: (2025)
by: Blake, Nathan, et al.
Published: (2025)
3D ReX: Causal Explanations in 3D Neuroimaging Classification
by: Navaratnarajah, Melane, et al.
Published: (2025)
by: Navaratnarajah, Melane, et al.
Published: (2025)
From Outcome-Based to Language-Based Preferences
by: Capraro, Valerio, et al.
Published: (2022)
by: Capraro, Valerio, et al.
Published: (2022)
GPT Deciphering Fedspeak: Quantifying Dissent Among Hawks and Doves
by: Peskoff, Denis, et al.
Published: (2024)
by: Peskoff, Denis, et al.
Published: (2024)
HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
by: Yang, Langqi, et al.
Published: (2025)
by: Yang, Langqi, et al.
Published: (2025)
Self-HarmLLM: Can Large Language Model Harm Itself?
by: Kim, Heehwan, et al.
Published: (2025)
by: Kim, Heehwan, et al.
Published: (2025)
When Harmful Content Gets Camouflaged: Unveiling Perception Failure of LVLMs with CamHarmTI
by: Li, Yanhui, et al.
Published: (2025)
by: Li, Yanhui, et al.
Published: (2025)
HarmfulSkillBench: How Do Harmful Skills Weaponize Your Agents?
by: Jiang, Yukun, et al.
Published: (2026)
by: Jiang, Yukun, et al.
Published: (2026)
Harm or Humor: A Multimodal, Multilingual Benchmark for Overt and Covert Harmful Humor
by: Sharshar, Ahmed, et al.
Published: (2026)
by: Sharshar, Ahmed, et al.
Published: (2026)
The Geometry of Harmfulness in LLMs through Subconcept Probing
by: Shah, McNair, et al.
Published: (2025)
by: Shah, McNair, et al.
Published: (2025)
HarmTransform: Transforming Explicit Harmful Queries into Stealthy via Multi-Agent Debate
by: Zhu, Shenzhe
Published: (2025)
by: Zhu, Shenzhe
Published: (2025)
Echoes of AI Harms: A Human-LLM Synergistic Framework for Bias-Driven Harm Anticipation
by: Tantalaki, Nicoleta, et al.
Published: (2025)
by: Tantalaki, Nicoleta, et al.
Published: (2025)
Booster: Tackling Harmful Fine-tuning for Large Language Models via Attenuating Harmful Perturbation
by: Huang, Tiansheng, et al.
Published: (2024)
by: Huang, Tiansheng, et al.
Published: (2024)
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
by: Andriushchenko, Maksym, et al.
Published: (2024)
by: Andriushchenko, Maksym, et al.
Published: (2024)
`For Argument's Sake, Show Me How to Harm Myself!': Jailbreaking LLMs in Suicide and Self-Harm Contexts
by: Schoene, Annika M, et al.
Published: (2025)
by: Schoene, Annika M, et al.
Published: (2025)
PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
by: Li, Jing-Jing, et al.
Published: (2026)
by: Li, Jing-Jing, et al.
Published: (2026)
SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
by: Pandey, Punya Syon, et al.
Published: (2025)
by: Pandey, Punya Syon, et al.
Published: (2025)
Why Do Large Language Models Generate Harmful Content?
by: Ganguli, Rajesh, et al.
Published: (2026)
by: Ganguli, Rajesh, et al.
Published: (2026)
Agentic AI Frameworks: Architectures, Protocols, and Design Challenges
by: Derouiche, Hana, et al.
Published: (2025)
by: Derouiche, Hana, et al.
Published: (2025)
Evaluating Language Models for Harmful Manipulation
by: Akbulut, Canfer, et al.
Published: (2026)
by: Akbulut, Canfer, et al.
Published: (2026)
Similar Items
-
Explaining Image Classifiers
by: Chockler, Hana, et al.
Published: (2024) -
Clustered Policy Decision Ranking
by: Levin, Mark, et al.
Published: (2023) -
Defining and Quantifying Creative Behavior in Popular Image Generators
by: Ramaswamy, Aditi, et al.
Published: (2025) -
Nondeterministic Causal Models
by: Beckers, Sander
Published: (2024) -
Large Language Models as Nondeterministic Causal Models
by: Beckers, Sander
Published: (2025)