HarmLevelBench: Evaluating Harm-Level Compliance and the Impact of Quantization on Model Alignment
Fuente:
arXiv
Guardado en:
| Autores principales: | Belkhiter, Yannis, Zizzo, Giulio, Maffeis, Sergio |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Breaking MCP with Function Hijacking Attacks: Novel Threats for Function Calling and Agentic Models
por: Belkhiter, Yannis, et al.
Publicado: (2026)
por: Belkhiter, Yannis, et al.
Publicado: (2026)
Elevating Defenses: Bridging Adversarial Training and Watermarking for Model Resilience
por: Thakkar, Janvi, et al.
Publicado: (2023)
por: Thakkar, Janvi, et al.
Publicado: (2023)
Segment-Level Coherence for Robust Harmful Intent Probing in LLMs
por: He, Xuanli, et al.
Publicado: (2026)
por: He, Xuanli, et al.
Publicado: (2026)
Blue Teaming Function-Calling Agents
por: Dolcetti, Greta, et al.
Publicado: (2026)
por: Dolcetti, Greta, et al.
Publicado: (2026)
Deep Research Brings Deeper Harm
por: Chen, Shuo, et al.
Publicado: (2025)
por: Chen, Shuo, et al.
Publicado: (2025)
ChineseHarm-Bench: A Chinese Harmful Content Detection Benchmark
por: Liu, Kangwei, et al.
Publicado: (2025)
por: Liu, Kangwei, et al.
Publicado: (2025)
Different Paths to Harmful Compliance: Behavioral Side Effects and Mechanistic Divergence Across LLM Jailbreaks
por: Kabir, Md Rysul, et al.
Publicado: (2026)
por: Kabir, Md Rysul, et al.
Publicado: (2026)
Jailbreaking Commercial Black-Box LLMs with Explicitly Harmful Prompts
por: Zhang, Chiyu, et al.
Publicado: (2025)
por: Zhang, Chiyu, et al.
Publicado: (2025)
CTRAP: Embedding Collapse Trap to Safeguard Large Language Models from Harmful Fine-Tuning
por: Yi, Biao, et al.
Publicado: (2025)
por: Yi, Biao, et al.
Publicado: (2025)
Expected Harm: Rethinking Safety Evaluation of (Mis)Aligned LLMs
por: Chen, Yen-Shan, et al.
Publicado: (2026)
por: Chen, Yen-Shan, et al.
Publicado: (2026)
Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models
por: An, Bang, et al.
Publicado: (2024)
por: An, Bang, et al.
Publicado: (2024)
HarmRLVR: Weaponizing Verifiable Rewards for Harmful LLM Alignment
por: Liu, Yuexiao, et al.
Publicado: (2025)
por: Liu, Yuexiao, et al.
Publicado: (2025)
Helpful or Harmful? Exploring the Efficacy of Large Language Models for Online Grooming Prevention
por: Prosser, Ellie, et al.
Publicado: (2024)
por: Prosser, Ellie, et al.
Publicado: (2024)
Safety Anchor: Defending Harmful Fine-tuning via Geometric Bottlenecks
por: Lu, Guoxin, et al.
Publicado: (2026)
por: Lu, Guoxin, et al.
Publicado: (2026)
RealHarm: A Collection of Real-World Language Model Application Failures
por: Jeune, Pierre Le, et al.
Publicado: (2025)
por: Jeune, Pierre Le, et al.
Publicado: (2025)
Latent Fusion Jailbreak: Blending Harmful and Harmless Representations to Elicit Unsafe LLM Outputs
por: Xing, Wenpeng, et al.
Publicado: (2025)
por: Xing, Wenpeng, et al.
Publicado: (2025)
Token-Level Privacy in Large Language Models
por: Harel, Re'em, et al.
Publicado: (2025)
por: Harel, Re'em, et al.
Publicado: (2025)
HarmfulSkillBench: How Do Harmful Skills Weaponize Your Agents?
por: Jiang, Yukun, et al.
Publicado: (2026)
por: Jiang, Yukun, et al.
Publicado: (2026)
GRAID: Synthetic Data Generation with Geometric Constraints and Multi-Agentic Reflection for Harmful Content Detection
por: Rad, Melissa Kazemi, et al.
Publicado: (2025)
por: Rad, Melissa Kazemi, et al.
Publicado: (2025)
When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Computer-Use Agents
por: Jones, Jaylen, et al.
Publicado: (2026)
por: Jones, Jaylen, et al.
Publicado: (2026)
Virus: Harmful Fine-tuning Attack for Large Language Models Bypassing Guardrail Moderation
por: Huang, Tiansheng, et al.
Publicado: (2025)
por: Huang, Tiansheng, et al.
Publicado: (2025)
A Generative Approach to LLM Harmfulness Mitigation with Red Flag Tokens
por: Dobre, David, et al.
Publicado: (2025)
por: Dobre, David, et al.
Publicado: (2025)
PrivLM-Bench: A Multi-level Privacy Evaluation Benchmark for Language Models
por: Li, Haoran, et al.
Publicado: (2023)
por: Li, Haoran, et al.
Publicado: (2023)
Step-Tagging: Toward controlling the generation of Language Reasoning Models through step monitoring
por: Belkhiter, Yannis, et al.
Publicado: (2025)
por: Belkhiter, Yannis, et al.
Publicado: (2025)
TRACES: Tagging Reasoning Steps for Adaptive Cost-Efficient Early-Stopping
por: Belkhiter, Yannis, et al.
Publicado: (2026)
por: Belkhiter, Yannis, et al.
Publicado: (2026)
Root Defence Strategies: Ensuring Safety of LLM at the Decoding Level
por: Zeng, Xinyi, et al.
Publicado: (2024)
por: Zeng, Xinyi, et al.
Publicado: (2024)
Eliciting Harmful Capabilities by Fine-Tuning On Safeguarded Outputs
por: Kaunismaa, Jackson, et al.
Publicado: (2026)
por: Kaunismaa, Jackson, et al.
Publicado: (2026)
GuidedBench: Measuring and Mitigating the Evaluation Discrepancies of In-the-wild LLM Jailbreak Methods
por: Huang, Ruixuan, et al.
Publicado: (2025)
por: Huang, Ruixuan, et al.
Publicado: (2025)
Towards Assuring EU AI Act Compliance and Adversarial Robustness of LLMs
por: Momcilovic, Tomas Bueno, et al.
Publicado: (2024)
por: Momcilovic, Tomas Bueno, et al.
Publicado: (2024)
Where Do Backdoors Live? A Component-Level Analysis of Backdoor Propagation in Speech Language Models
por: Fortier, Alexandrine, et al.
Publicado: (2025)
por: Fortier, Alexandrine, et al.
Publicado: (2025)
EmMark: Robust Watermarks for IP Protection of Embedded Quantized Large Language Models
por: Zhang, Ruisi, et al.
Publicado: (2024)
por: Zhang, Ruisi, et al.
Publicado: (2024)
ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models
por: Zhao, Yunhan, et al.
Publicado: (2026)
por: Zhao, Yunhan, et al.
Publicado: (2026)
Developing Assurance Cases for Adversarial Robustness and Regulatory Compliance in LLMs
por: Momcilovic, Tomas Bueno, et al.
Publicado: (2024)
por: Momcilovic, Tomas Bueno, et al.
Publicado: (2024)
Safety and Security Analysis of Large Language Models: Benchmarking Risk Profile and Harm Potential
por: Akiri, Charankumar, et al.
Publicado: (2025)
por: Akiri, Charankumar, et al.
Publicado: (2025)
Self and Cross-Model Distillation for LLMs: Effective Methods for Refusal Pattern Alignment
por: Li, Jie, et al.
Publicado: (2024)
por: Li, Jie, et al.
Publicado: (2024)
Mind the Privacy Unit! User-Level Differential Privacy for Language Model Fine-Tuning
por: Chua, Lynn, et al.
Publicado: (2024)
por: Chua, Lynn, et al.
Publicado: (2024)
DocMIA: Document-Level Membership Inference Attacks against DocVQA Models
por: Nguyen, Khanh, et al.
Publicado: (2025)
por: Nguyen, Khanh, et al.
Publicado: (2025)
EnchTable: Unified Safety Alignment Transfer in Fine-tuned Large Language Models
por: Wu, Jialin, et al.
Publicado: (2025)
por: Wu, Jialin, et al.
Publicado: (2025)
On the Generation and Mitigation of Harmful Geometry in Image-to-3D Models
por: Liu, Yule, et al.
Publicado: (2026)
por: Liu, Yule, et al.
Publicado: (2026)
Verifiability and Privacy in Federated Learning through Context-Hiding Multi-Key Homomorphic Authenticators
por: Bottoni, Simone, et al.
Publicado: (2025)
por: Bottoni, Simone, et al.
Publicado: (2025)
Ejemplares similares
-
Breaking MCP with Function Hijacking Attacks: Novel Threats for Function Calling and Agentic Models
por: Belkhiter, Yannis, et al.
Publicado: (2026) -
Elevating Defenses: Bridging Adversarial Training and Watermarking for Model Resilience
por: Thakkar, Janvi, et al.
Publicado: (2023) -
Segment-Level Coherence for Robust Harmful Intent Probing in LLMs
por: He, Xuanli, et al.
Publicado: (2026) -
Blue Teaming Function-Calling Agents
por: Dolcetti, Greta, et al.
Publicado: (2026) -
Deep Research Brings Deeper Harm
por: Chen, Shuo, et al.
Publicado: (2025)