Expected Harm: Rethinking Safety Evaluation of (Mis)Aligned LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Yen-Shan, Tam, Zhi Rui, Wu, Cheng-Kuang, Chen, Yun-Nung |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Eyes-on-Me: Scalable RAG Poisoning through Transferable Attention-Steering Attractors
by: Chen, Yen-Shan, et al.
Published: (2025)
by: Chen, Yen-Shan, et al.
Published: (2025)
None of the Above, Less of the Right: Parallel Patterns between Humans and LLMs on Multi-Choice Questions Answering
by: Tam, Zhi Rui, et al.
Published: (2025)
by: Tam, Zhi Rui, et al.
Published: (2025)
Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models
by: An, Bang, et al.
Published: (2024)
by: An, Bang, et al.
Published: (2024)
TraceSafe: A Systematic Assessment of LLM Guardrails on Multi-Step Tool-Calling Trajectories
by: Chen, Yen-Shan, et al.
Published: (2026)
by: Chen, Yen-Shan, et al.
Published: (2026)
Gaming the Metric, Not the Harm: Certifying Safety Audits against Strategic Platform Manipulation
by: Burnat, Florian A. D., et al.
Published: (2026)
by: Burnat, Florian A. D., et al.
Published: (2026)
The Art of (Mis)alignment: How Fine-Tuning Methods Effectively Misalign and Realign LLMs in Post-Training
by: Zhang, Rui, et al.
Published: (2026)
by: Zhang, Rui, et al.
Published: (2026)
SafeCOMM: A Study on Safety Degradation in Fine-Tuned Telecom Large Language Models
by: Djuhera, Aladin, et al.
Published: (2025)
by: Djuhera, Aladin, et al.
Published: (2025)
Medical Malice: A Dataset for Context-Aware Safety in Healthcare LLMs
by: D'addario, Andrew Maranhão Ventura
Published: (2025)
by: D'addario, Andrew Maranhão Ventura
Published: (2025)
Superficial Safety Alignment Hypothesis
by: Li, Jianwei, et al.
Published: (2024)
by: Li, Jianwei, et al.
Published: (2024)
ConVerse: Benchmarking Contextual Safety in Agent-to-Agent Conversations
by: Gomaa, Amr, et al.
Published: (2025)
by: Gomaa, Amr, et al.
Published: (2025)
RealHarm: A Collection of Real-World Language Model Application Failures
by: Jeune, Pierre Le, et al.
Published: (2025)
by: Jeune, Pierre Le, et al.
Published: (2025)
Safety and Security Analysis of Large Language Models: Benchmarking Risk Profile and Harm Potential
by: Akiri, Charankumar, et al.
Published: (2025)
by: Akiri, Charankumar, et al.
Published: (2025)
Jailbreaking Commercial Black-Box LLMs with Explicitly Harmful Prompts
by: Zhang, Chiyu, et al.
Published: (2025)
by: Zhang, Chiyu, et al.
Published: (2025)
Digger: Detecting Copyright Content Mis-usage in Large Language Model Training
by: Li, Haodong, et al.
Published: (2024)
by: Li, Haodong, et al.
Published: (2024)
Machine Unlearning Meets Adversarial Robustness via Constrained Interventions on LLMs
by: Rezkellah, Fatmazohra, et al.
Published: (2025)
by: Rezkellah, Fatmazohra, et al.
Published: (2025)
Let's Measure the Elephant in the Room: Facilitating Personalized Automated Analysis of Privacy Policies at Scale
by: Zhao, Rui, et al.
Published: (2025)
by: Zhao, Rui, et al.
Published: (2025)
A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy
by: Wang, Huandong, et al.
Published: (2025)
by: Wang, Huandong, et al.
Published: (2025)
k-SemStamp: A Clustering-Based Semantic Watermark for Detection of Machine-Generated Text
by: Hou, Abe Bohan, et al.
Published: (2024)
by: Hou, Abe Bohan, et al.
Published: (2024)
SimMark: A Robust Sentence-Level Similarity-Based Watermarking Algorithm for Large Language Models
by: Dabiriaghdam, Amirhossein, et al.
Published: (2025)
by: Dabiriaghdam, Amirhossein, et al.
Published: (2025)
BadFair: Backdoored Fairness Attacks with Group-conditioned Triggers
by: Xue, Jiaqi, et al.
Published: (2024)
by: Xue, Jiaqi, et al.
Published: (2024)
In the Name of Fairness: Assessing the Bias in Clinical Record De-identification
by: Xiao, Yuxin, et al.
Published: (2023)
by: Xiao, Yuxin, et al.
Published: (2023)
LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet
by: Li, Nathaniel, et al.
Published: (2024)
by: Li, Nathaniel, et al.
Published: (2024)
GRAID: Synthetic Data Generation with Geometric Constraints and Multi-Agentic Reflection for Harmful Content Detection
by: Rad, Melissa Kazemi, et al.
Published: (2025)
by: Rad, Melissa Kazemi, et al.
Published: (2025)
HarmLevelBench: Evaluating Harm-Level Compliance and the Impact of Quantization on Model Alignment
by: Belkhiter, Yannis, et al.
Published: (2024)
by: Belkhiter, Yannis, et al.
Published: (2024)
How Susceptible are Large Language Models to Ideological Manipulation?
by: Chen, Kai, et al.
Published: (2024)
by: Chen, Kai, et al.
Published: (2024)
Segment-Level Coherence for Robust Harmful Intent Probing in LLMs
by: He, Xuanli, et al.
Published: (2026)
by: He, Xuanli, et al.
Published: (2026)
What Makes an Evaluation Useful? Common Pitfalls and Best Practices
by: Gekker, Gil, et al.
Published: (2025)
by: Gekker, Gil, et al.
Published: (2025)
Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
by: Zhang, Andy K., et al.
Published: (2024)
by: Zhang, Andy K., et al.
Published: (2024)
Deep Research Brings Deeper Harm
by: Chen, Shuo, et al.
Published: (2025)
by: Chen, Shuo, et al.
Published: (2025)
Can a large language model be a gaslighter?
by: Li, Wei, et al.
Published: (2024)
by: Li, Wei, et al.
Published: (2024)
LLM Platform Security: Applying a Systematic Evaluation Framework to OpenAI's ChatGPT Plugins
by: Iqbal, Umar, et al.
Published: (2023)
by: Iqbal, Umar, et al.
Published: (2023)
Phare: A Safety Probe for Large Language Models
by: Jeune, Pierre Le, et al.
Published: (2025)
by: Jeune, Pierre Le, et al.
Published: (2025)
An Evaluation of Chat Safety Moderations in Roblox
by: Kaushik, Priya, et al.
Published: (2026)
by: Kaushik, Priya, et al.
Published: (2026)
Cross-Task Defense: Instruction-Tuning LLMs for Content Safety
by: Fu, Yu, et al.
Published: (2024)
by: Fu, Yu, et al.
Published: (2024)
An In-Depth Investigation of Data Collection in LLM App Ecosystems
by: Wu, Yuhao, et al.
Published: (2024)
by: Wu, Yuhao, et al.
Published: (2024)
IsolateGPT: An Execution Isolation Architecture for LLM-Based Agentic Systems
by: Wu, Yuhao, et al.
Published: (2024)
by: Wu, Yuhao, et al.
Published: (2024)
ROK-FORTRESS: Measuring the Effect of Geopolitical Transcreation for National Security and Public Safety
by: Lee, Michael S., et al.
Published: (2026)
by: Lee, Michael S., et al.
Published: (2026)
How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States
by: Zhou, Zhenhong, et al.
Published: (2024)
by: Zhou, Zhenhong, et al.
Published: (2024)
Rethinking Anonymity Claims in Synthetic Data Generation: A Model-Centric Privacy Attack Perspective
by: Ganev, Georgi, et al.
Published: (2026)
by: Ganev, Georgi, et al.
Published: (2026)
Clio: Privacy-Preserving Insights into Real-World AI Use
by: Tamkin, Alex, et al.
Published: (2024)
by: Tamkin, Alex, et al.
Published: (2024)
Similar Items
-
Eyes-on-Me: Scalable RAG Poisoning through Transferable Attention-Steering Attractors
by: Chen, Yen-Shan, et al.
Published: (2025) -
None of the Above, Less of the Right: Parallel Patterns between Humans and LLMs on Multi-Choice Questions Answering
by: Tam, Zhi Rui, et al.
Published: (2025) -
Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models
by: An, Bang, et al.
Published: (2024) -
TraceSafe: A Systematic Assessment of LLM Guardrails on Multi-Step Tool-Calling Trajectories
by: Chen, Yen-Shan, et al.
Published: (2026) -
Gaming the Metric, Not the Harm: Certifying Safety Audits against Strategic Platform Manipulation
by: Burnat, Florian A. D., et al.
Published: (2026)