Evaluation without Generation: Non-Generative Assessment of Harmful Model Specialization with Applications to CSAM
Fuente:
arXiv
Saved in:
| Main Authors: | Suriyakumar, Vinith M., Sekhari, Ayush, Stempfle, Lena, Wang, Robertson, Simpson, Michael, Portnoff, Rebecca, Ghassemi, Marzyeh, Wilson, Ashia C. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UCD: Unlearning in LLMs via Contrastive Decoding
by: Suriyakumar, Vinith M., et al.
Published: (2025)
by: Suriyakumar, Vinith M., et al.
Published: (2025)
Algorithmic Pluralism: A Structural Approach To Equal Opportunity
by: Jain, Shomik, et al.
Published: (2023)
by: Jain, Shomik, et al.
Published: (2023)
When Style Breaks Safety: Defending LLMs Against Superficial Style Alignment
by: Xiao, Yuxin, et al.
Published: (2025)
by: Xiao, Yuxin, et al.
Published: (2025)
Unstable Unlearning: The Hidden Risk of Concept Resurgence in Diffusion Models
by: Suriyakumar, Vinith M., et al.
Published: (2024)
by: Suriyakumar, Vinith M., et al.
Published: (2024)
AI Generated Child Sexual Abuse Material -- What's the Harm?
by: Ciardha, Caoilte Ó, et al.
Published: (2025)
by: Ciardha, Caoilte Ó, et al.
Published: (2025)
Layered Unlearning for Adversarial Relearning
by: Qian, Timothy, et al.
Published: (2025)
by: Qian, Timothy, et al.
Published: (2025)
Learning the Wrong Lessons: Syntactic-Domain Spurious Correlations in Language Models
by: Shaib, Chantal, et al.
Published: (2025)
by: Shaib, Chantal, et al.
Published: (2025)
Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple Interactions
by: Chan, Yik Siu, et al.
Published: (2025)
by: Chan, Yik Siu, et al.
Published: (2025)
An Investigation of Memorization Risk in Healthcare Foundation Models
by: Tonekaboni, Sana, et al.
Published: (2025)
by: Tonekaboni, Sana, et al.
Published: (2025)
Position: AI Evaluations Should be Grounded on a Theory of Capability
by: Jo, Nathanael, et al.
Published: (2025)
by: Jo, Nathanael, et al.
Published: (2025)
Bias Delayed is Bias Denied? Assessing the Effect of Reporting Delays on Disparity Assessments
by: Gosciak, Jennah, et al.
Published: (2025)
by: Gosciak, Jennah, et al.
Published: (2025)
Generative AI in Medicine
by: Shanmugam, Divya, et al.
Published: (2024)
by: Shanmugam, Divya, et al.
Published: (2024)
What's in a Query: Polarity-Aware Distribution-Based Fair Ranking
by: Balagopalan, Aparna, et al.
Published: (2025)
by: Balagopalan, Aparna, et al.
Published: (2025)
As an AI Language Model, "Yes I Would Recommend Calling the Police": Norm Inconsistency in LLM Decision-Making
by: Jain, Shomik, et al.
Published: (2024)
by: Jain, Shomik, et al.
Published: (2024)
Allocation Multiplicity: Evaluating the Promises of the Rashomon Set
by: Jain, Shomik, et al.
Published: (2025)
by: Jain, Shomik, et al.
Published: (2025)
In the Name of Fairness: Assessing the Bias in Clinical Record De-identification
by: Xiao, Yuxin, et al.
Published: (2023)
by: Xiao, Yuxin, et al.
Published: (2023)
Unveiling AI's Threats to Child Protection: Regulatory efforts to Criminalize AI-Generated CSAM and Emerging Children's Rights Violations
by: Kokolaki, Emmanouela, et al.
Published: (2025)
by: Kokolaki, Emmanouela, et al.
Published: (2025)
Just in Plain Sight: Unveiling CSAM Distribution Campaigns on the Clear Web
by: Lykousas, Nikolaos, et al.
Published: (2025)
by: Lykousas, Nikolaos, et al.
Published: (2025)
Scarce Resource Allocations That Rely On Machine Learning Should Be Randomized
by: Jain, Shomik, et al.
Published: (2024)
by: Jain, Shomik, et al.
Published: (2024)
Data Debiasing with Datamodels (D3M): Improving Subgroup Robustness via Data Selection
by: Jain, Saachi, et al.
Published: (2024)
by: Jain, Saachi, et al.
Published: (2024)
DiffusionWorldViewer: Exposing and Broadening the Worldview Reflected by Generative Text-to-Image Models
by: De Simone, Zoe, et al.
Published: (2023)
by: De Simone, Zoe, et al.
Published: (2023)
Automating Transparency Mechanisms in the Judicial System Using LLMs: Opportunities and Challenges
by: Shastri, Ishana, et al.
Published: (2024)
by: Shastri, Ishana, et al.
Published: (2024)
The Gaussian Mixing Mechanism: Renyi Differential Privacy via Gaussian Sketches
by: Lev, Omri, et al.
Published: (2025)
by: Lev, Omri, et al.
Published: (2025)
Identifying Implicit Social Biases in Vision-Language Models
by: Hamidieh, Kimia, et al.
Published: (2024)
by: Hamidieh, Kimia, et al.
Published: (2024)
Evaluation eines Lehramtsmasterstudiengangs mit dem Profil Quereinstieg im Fach Physik
by: Ghassemi, Novid
Published: (2024)
by: Ghassemi, Novid
Published: (2024)
Hidden Poison: Machine Unlearning Enables Camouflaged Poisoning Attacks
by: Di, Jimmy Z., et al.
Published: (2022)
by: Di, Jimmy Z., et al.
Published: (2022)
Homogeneous Algorithms Can Reduce Competition in Personalized Pricing
by: Jo, Nathanael, et al.
Published: (2025)
by: Jo, Nathanael, et al.
Published: (2025)
Machine Unlearning Fails to Remove Data Poisoning Attacks
by: Pawelczyk, Martin, et al.
Published: (2024)
by: Pawelczyk, Martin, et al.
Published: (2024)
Task-Dependent Evaluation of LLM Output Homogenization: A Taxonomy-Guided Framework
by: Jain, Shomik, et al.
Published: (2025)
by: Jain, Shomik, et al.
Published: (2025)
Towards a Harms Taxonomy of AI Likeness Generation
by: Bariach, Ben, et al.
Published: (2024)
by: Bariach, Ben, et al.
Published: (2024)
Do Generative AI Models Output Harm while Representing Non-Western Cultures: Evidence from A Community-Centered Approach
by: Ghosh, Sourojit, et al.
Published: (2024)
by: Ghosh, Sourojit, et al.
Published: (2024)
Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models
by: An, Bang, et al.
Published: (2024)
by: An, Bang, et al.
Published: (2024)
Hedging and Non-Affirmation: Quantifying LLM Alignment on Questions of Human Rights
by: Javed, Rafiya, et al.
Published: (2025)
by: Javed, Rafiya, et al.
Published: (2025)
Designing Incident Reporting Systems for Harms from General-Purpose AI
by: Wei, Kevin, et al.
Published: (2025)
by: Wei, Kevin, et al.
Published: (2025)
Safety and Security Analysis of Large Language Models: Benchmarking Risk Profile and Harm Potential
by: Akiri, Charankumar, et al.
Published: (2025)
by: Akiri, Charankumar, et al.
Published: (2025)
Perpetuating Misogyny with Generative AI: How Model Personalization Normalizes Gendered Harm
by: Wagner, Laura, et al.
Published: (2025)
by: Wagner, Laura, et al.
Published: (2025)
Evaluating Language Models for Harmful Manipulation
by: Akbulut, Canfer, et al.
Published: (2026)
by: Akbulut, Canfer, et al.
Published: (2026)
Machine-arranged Interactions Improve Institutional Belonging and Cohesion
by: Ghassemi, Mohammad M., et al.
Published: (2024)
by: Ghassemi, Mohammad M., et al.
Published: (2024)
A Mechanism-Based Approach to Mitigating Harms from Persuasive Generative AI
by: El-Sayed, Seliem, et al.
Published: (2024)
by: El-Sayed, Seliem, et al.
Published: (2024)
Crafting Tomorrow's Evaluations: Assessment Design Strategies in the Era of Generative AI
by: Kadel, Rajan, et al.
Published: (2024)
by: Kadel, Rajan, et al.
Published: (2024)
Similar Items
-
UCD: Unlearning in LLMs via Contrastive Decoding
by: Suriyakumar, Vinith M., et al.
Published: (2025) -
Algorithmic Pluralism: A Structural Approach To Equal Opportunity
by: Jain, Shomik, et al.
Published: (2023) -
When Style Breaks Safety: Defending LLMs Against Superficial Style Alignment
by: Xiao, Yuxin, et al.
Published: (2025) -
Unstable Unlearning: The Hidden Risk of Concept Resurgence in Diffusion Models
by: Suriyakumar, Vinith M., et al.
Published: (2024) -
AI Generated Child Sexual Abuse Material -- What's the Harm?
by: Ciardha, Caoilte Ó, et al.
Published: (2025)