Safety and Security Analysis of Large Language Models: Benchmarking Risk Profile and Harm Potential
Fuente:
arXiv
Saved in:
| Main Authors: | Akiri, Charankumar, Simpson, Harrison, Aryal, Kshitiz, Khanna, Aarav, Gupta, Maanak |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Explainability-Informed Targeted Malware Misclassification
by: Card, Quincy, et al.
Published: (2024)
by: Card, Quincy, et al.
Published: (2024)
Explainability-Guided Adversarial Attacks on Transformer-Based Malware Detectors Using Control Flow Graphs
by: Wheeler, Andrew, et al.
Published: (2026)
by: Wheeler, Andrew, et al.
Published: (2026)
RAG-targeted Adversarial Attack on LLM-based Threat Detection and Mitigation Framework
by: Ikbarieh, Seif, et al.
Published: (2025)
by: Ikbarieh, Seif, et al.
Published: (2025)
Explainable Deep Learning Models for Dynamic and Online Malware Classification
by: Card, Quincy, et al.
Published: (2024)
by: Card, Quincy, et al.
Published: (2024)
Explainability Guided Adversarial Evasion Attacks on Malware Detectors
by: Aryal, Kshitiz, et al.
Published: (2024)
by: Aryal, Kshitiz, et al.
Published: (2024)
Intra-Section Code Cave Injection for Adversarial Evasion Attacks on Windows PE Malware File
by: Aryal, Kshitiz, et al.
Published: (2024)
by: Aryal, Kshitiz, et al.
Published: (2024)
SoK: Leveraging Transformers for Malware Analysis
by: Kunwar, Pradip, et al.
Published: (2024)
by: Kunwar, Pradip, et al.
Published: (2024)
A Survey of Agentic AI and Cybersecurity: Challenges, Opportunities and Use-case Prototypes
by: Lazer, Sahaya Jestus, et al.
Published: (2026)
by: Lazer, Sahaya Jestus, et al.
Published: (2026)
Identifying Security Risks in NFT Platforms
by: Gupta, Yash, et al.
Published: (2022)
by: Gupta, Yash, et al.
Published: (2022)
Evaluating the Critical Risks of Amazon's Nova Premier under the Frontier Model Safety Framework
by: Krishna, Satyapriya, et al.
Published: (2025)
by: Krishna, Satyapriya, et al.
Published: (2025)
Large Language Models as a (Bad) Security Norm in the Context of Regulation and Compliance
by: Ludvigsen, Kaspar Rosager
Published: (2025)
by: Ludvigsen, Kaspar Rosager
Published: (2025)
Benchmarking and Understanding Safety Risks in AI Character Platforms
by: Wei, Yiluo, et al.
Published: (2025)
by: Wei, Yiluo, et al.
Published: (2025)
An Investigation into Misuse of Java Security APIs by Large Language Models
by: Mousavi, Zahra, et al.
Published: (2024)
by: Mousavi, Zahra, et al.
Published: (2024)
The Incoherency Risk in the EU's New Cyber Security Policies
by: Ruohonen, Jukka
Published: (2024)
by: Ruohonen, Jukka
Published: (2024)
Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models
by: An, Bang, et al.
Published: (2024)
by: An, Bang, et al.
Published: (2024)
Gaming the Metric, Not the Harm: Certifying Safety Audits against Strategic Platform Manipulation
by: Burnat, Florian A. D., et al.
Published: (2026)
by: Burnat, Florian A. D., et al.
Published: (2026)
Securing the Web: Analysis of HTTP Security Headers in Popular Global Websites
by: Kishnani, Urvashi, et al.
Published: (2024)
by: Kishnani, Urvashi, et al.
Published: (2024)
Backchaining Loss of Control Mitigations from Mission-Specific Benchmarks in National Security
by: Pistillo, Matteo, et al.
Published: (2026)
by: Pistillo, Matteo, et al.
Published: (2026)
Expected Harm: Rethinking Safety Evaluation of (Mis)Aligned LLMs
by: Chen, Yen-Shan, et al.
Published: (2026)
by: Chen, Yen-Shan, et al.
Published: (2026)
Leveraging Large Language Models for Preliminary Security Risk Analysis: A Mission-Critical Case Study
by: Esposito, Matteo, et al.
Published: (2024)
by: Esposito, Matteo, et al.
Published: (2024)
A Security Assessment tool for Quantum Threat Analysis
by: Halak, Basel, et al.
Published: (2024)
by: Halak, Basel, et al.
Published: (2024)
A Lightweight Edge-CNN-Transformer Model for Detecting Coordinated Cyber and Digital Twin Attacks in Cooperative Smart Farming
by: Praharaj, Lopamudra, et al.
Published: (2024)
by: Praharaj, Lopamudra, et al.
Published: (2024)
Phare: A Safety Probe for Large Language Models
by: Jeune, Pierre Le, et al.
Published: (2025)
by: Jeune, Pierre Le, et al.
Published: (2025)
Generative AI Misuse Potential in Cyber Security Education: A Case Study of a UK Degree Program
by: Shepherd, Carlton
Published: (2025)
by: Shepherd, Carlton
Published: (2025)
ConVerse: Benchmarking Contextual Safety in Agent-to-Agent Conversations
by: Gomaa, Amr, et al.
Published: (2025)
by: Gomaa, Amr, et al.
Published: (2025)
What's Privacy Good for? Measuring Privacy as a Shield from Harms due to Personal Data Use
by: Gajavalli, Sri Harsha, et al.
Published: (2025)
by: Gajavalli, Sri Harsha, et al.
Published: (2025)
$\texttt{ModSCAN}$: Measuring Stereotypical Bias in Large Vision-Language Models from Vision and Language Modalities
by: Jiang, Yukun, et al.
Published: (2024)
by: Jiang, Yukun, et al.
Published: (2024)
A Proposal for Evaluating the Operational Risk for ChatBots based on Large Language Models
by: Pinacho-Davidson, Pedro, et al.
Published: (2025)
by: Pinacho-Davidson, Pedro, et al.
Published: (2025)
RealHarm: A Collection of Real-World Language Model Application Failures
by: Jeune, Pierre Le, et al.
Published: (2025)
by: Jeune, Pierre Le, et al.
Published: (2025)
AI Risk-Management Standards Profile for General-Purpose AI (GPAI) and Foundation Models
by: Barrett, Anthony M., et al.
Published: (2025)
by: Barrett, Anthony M., et al.
Published: (2025)
Dual-Technique Privacy & Security Analysis for E-Commerce Websites Through Automated and Manual Implementation
by: Kishnani, Urvashi, et al.
Published: (2024)
by: Kishnani, Urvashi, et al.
Published: (2024)
AI Safety vs. AI Security: Demystifying the Distinction and Boundaries
by: Lin, Zhiqiang, et al.
Published: (2025)
by: Lin, Zhiqiang, et al.
Published: (2025)
MM-AttacKG: A Multimodal Approach to Attack Graph Construction with Large Language Models
by: Zhang, Yongheng, et al.
Published: (2025)
by: Zhang, Yongheng, et al.
Published: (2025)
Modeling Behavioral Signals in Job Scams: A Human-Centered Security Study
by: Anagha, Goni, et al.
Published: (2026)
by: Anagha, Goni, et al.
Published: (2026)
From Cyber Security Incident Management to Cyber Security Crisis Management in the European Union
by: Ruohonen, Jukka, et al.
Published: (2025)
by: Ruohonen, Jukka, et al.
Published: (2025)
Risks and Compliance with the EU's Core Cyber Security Legislation
by: Ruohonen, Jukka, et al.
Published: (2025)
by: Ruohonen, Jukka, et al.
Published: (2025)
Technologies and Security Challenges in Metaverse
by: Dey, Krishno, et al.
Published: (2025)
by: Dey, Krishno, et al.
Published: (2025)
Security practices in AI development
by: Spelda, Petr, et al.
Published: (2025)
by: Spelda, Petr, et al.
Published: (2025)
Data Defenses Against Large Language Models
by: Agnew, William, et al.
Published: (2024)
by: Agnew, William, et al.
Published: (2024)
User Privacy and Large Language Models: An Analysis of Frontier Developers' Privacy Policies
by: King, Jennifer, et al.
Published: (2025)
by: King, Jennifer, et al.
Published: (2025)
Similar Items
-
Explainability-Informed Targeted Malware Misclassification
by: Card, Quincy, et al.
Published: (2024) -
Explainability-Guided Adversarial Attacks on Transformer-Based Malware Detectors Using Control Flow Graphs
by: Wheeler, Andrew, et al.
Published: (2026) -
RAG-targeted Adversarial Attack on LLM-based Threat Detection and Mitigation Framework
by: Ikbarieh, Seif, et al.
Published: (2025) -
Explainable Deep Learning Models for Dynamic and Online Malware Classification
by: Card, Quincy, et al.
Published: (2024) -
Explainability Guided Adversarial Evasion Attacks on Malware Detectors
by: Aryal, Kshitiz, et al.
Published: (2024)