aiXamine: Simplified LLM Safety and Security
Fuente:
arXiv
Salvato in:
| Autori principali: | Deniz, Fatih, Popovic, Dorde, Boshmaf, Yazan, Jeong, Euisuh, Ahmad, Minhaj, Chawla, Sanjay, Khalil, Issa |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DeBackdoor: A Deductive Framework for Detecting Backdoor Attacks on Deep Models with Limited Data
di: Popovic, Dorde, et al.
Pubblicazione: (2025)
di: Popovic, Dorde, et al.
Pubblicazione: (2025)
CallShield: Secure Caller Authentication over Real-Time Audio Channels
di: Rabh, Mouna, et al.
Pubblicazione: (2026)
di: Rabh, Mouna, et al.
Pubblicazione: (2026)
Swiss-Bench 003: Evaluating LLM Reliability and Adversarial Security for Swiss Regulatory Contexts
di: Uenal, Fatih
Pubblicazione: (2026)
di: Uenal, Fatih
Pubblicazione: (2026)
MANTIS: Detection of Zero-Day Malicious Domains Leveraging Low Reputed Hosting Infrastructure
di: Deniz, Fatih, et al.
Pubblicazione: (2025)
di: Deniz, Fatih, et al.
Pubblicazione: (2025)
Simplified and Secure MCP Gateways for Enterprise AI Integration
di: Brett, Ivo
Pubblicazione: (2025)
di: Brett, Ivo
Pubblicazione: (2025)
Explainable AI-based Intrusion Detection System for Industry 5.0: An Overview of the Literature, associated Challenges, the existing Solutions, and Potential Research Directions
di: Khan, Naseem, et al.
Pubblicazione: (2024)
di: Khan, Naseem, et al.
Pubblicazione: (2024)
Demo: SGCode: A Flexible Prompt-Optimizing System for Secure Generation of Code
di: Ton, Khiem, et al.
Pubblicazione: (2024)
di: Ton, Khiem, et al.
Pubblicazione: (2024)
Poison with Style: A Practical Poisoning Attack on Code Large Language Models
di: Tran, Khang, et al.
Pubblicazione: (2026)
di: Tran, Khang, et al.
Pubblicazione: (2026)
Safety Layers in Aligned Large Language Models: The Key to LLM Security
di: Li, Shen, et al.
Pubblicazione: (2024)
di: Li, Shen, et al.
Pubblicazione: (2024)
FlipLLM: Efficient Bit-Flip Attacks on Multimodal LLMs using Reinforcement Learning
di: Khalil, Khurram, et al.
Pubblicazione: (2025)
di: Khalil, Khurram, et al.
Pubblicazione: (2025)
Seven Security Challenges That Must be Solved in Cross-domain Multi-agent LLM Systems
di: Ko, Ronny, et al.
Pubblicazione: (2025)
di: Ko, Ronny, et al.
Pubblicazione: (2025)
Human Society-Inspired Approaches to Agentic AI Security: The 4C Framework
di: Abuadbba, Alsharif, et al.
Pubblicazione: (2026)
di: Abuadbba, Alsharif, et al.
Pubblicazione: (2026)
On the (In)Security of LLM App Stores
di: Hou, Xinyi, et al.
Pubblicazione: (2024)
di: Hou, Xinyi, et al.
Pubblicazione: (2024)
Proactively Detecting Threats: A Novel Approach Using LLMs
di: Chawla, Aniesh, et al.
Pubblicazione: (2026)
di: Chawla, Aniesh, et al.
Pubblicazione: (2026)
A Decompilation-Driven Framework for Malware Detection with Large Language Models
di: Chawla, Aniesh, et al.
Pubblicazione: (2026)
di: Chawla, Aniesh, et al.
Pubblicazione: (2026)
A Unified Evaluation of Learning-Based Similarity Techniques for Malware Detection
di: Prasad, Udbhav, et al.
Pubblicazione: (2026)
di: Prasad, Udbhav, et al.
Pubblicazione: (2026)
Cisco Integrated AI Security and Safety Framework Report
di: Chang, Amy, et al.
Pubblicazione: (2025)
di: Chang, Amy, et al.
Pubblicazione: (2025)
Measuring Safety Alignment Effects in Autonomous Security Agents
di: David, Isaac, et al.
Pubblicazione: (2026)
di: David, Isaac, et al.
Pubblicazione: (2026)
SoK: Towards Security and Safety of Edge AI
di: Wingarz, Tatjana, et al.
Pubblicazione: (2024)
di: Wingarz, Tatjana, et al.
Pubblicazione: (2024)
CyberLLMInstruct: A Pseudo-malicious Dataset Revealing Safety-performance Trade-offs in Cyber Security LLM Fine-tuning
di: ElZemity, Adel, et al.
Pubblicazione: (2025)
di: ElZemity, Adel, et al.
Pubblicazione: (2025)
Security and Privacy Issues and Solutions in Federated Learning for Digital Healthcare
di: Jeong, Hyejun, et al.
Pubblicazione: (2024)
di: Jeong, Hyejun, et al.
Pubblicazione: (2024)
EchoLeak: The First Real-World Zero-Click Prompt Injection Exploit in a Production LLM System
di: Reddy, Pavan, et al.
Pubblicazione: (2025)
di: Reddy, Pavan, et al.
Pubblicazione: (2025)
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning
di: Zhang, Yi, et al.
Pubblicazione: (2025)
di: Zhang, Yi, et al.
Pubblicazione: (2025)
A New Era in LLM Security: Exploring Security Concerns in Real-World LLM-based Systems
di: Wu, Fangzhou, et al.
Pubblicazione: (2024)
di: Wu, Fangzhou, et al.
Pubblicazione: (2024)
LLM-Safety Evaluations Lack Robustness
di: Beyer, Tim, et al.
Pubblicazione: (2025)
di: Beyer, Tim, et al.
Pubblicazione: (2025)
Systematization of Knowledge: Security and Safety in the Model Context Protocol Ecosystem
di: Gaire, Shiva, et al.
Pubblicazione: (2025)
di: Gaire, Shiva, et al.
Pubblicazione: (2025)
AI Risk Management Should Incorporate Both Safety and Security
di: Qi, Xiangyu, et al.
Pubblicazione: (2024)
di: Qi, Xiangyu, et al.
Pubblicazione: (2024)
TIPS: Threat Actor Informed Prioritization of Applications using SecEncoder
di: Bulut, Muhammed Fatih, et al.
Pubblicazione: (2024)
di: Bulut, Muhammed Fatih, et al.
Pubblicazione: (2024)
Differential Privacy-Driven Framework for Enhancing Heart Disease Prediction
di: Otoum, Yazan, et al.
Pubblicazione: (2025)
di: Otoum, Yazan, et al.
Pubblicazione: (2025)
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications
di: David, Isaac, et al.
Pubblicazione: (2026)
di: David, Isaac, et al.
Pubblicazione: (2026)
LLM Agents Should Employ Security Principles
di: Zhang, Kaiyuan, et al.
Pubblicazione: (2025)
di: Zhang, Kaiyuan, et al.
Pubblicazione: (2025)
A Framework for Formalizing LLM Agent Security
di: Siu, Vincent, et al.
Pubblicazione: (2026)
di: Siu, Vincent, et al.
Pubblicazione: (2026)
Blockchain Meets Adaptive Honeypots: A Trust-Aware Approach to Next-Gen IoT Security
di: Otoum, Yazan, et al.
Pubblicazione: (2025)
di: Otoum, Yazan, et al.
Pubblicazione: (2025)
Hybrid LLM-Enhanced Intrusion Detection for Zero-Day Threats in IoT Networks
di: Al-Hammouri, Mohammad F., et al.
Pubblicazione: (2025)
di: Al-Hammouri, Mohammad F., et al.
Pubblicazione: (2025)
NOIR: Privacy-Preserving Generation of Code with Open-Source LLMs
di: Nguyen, Khoa, et al.
Pubblicazione: (2026)
di: Nguyen, Khoa, et al.
Pubblicazione: (2026)
SAGE: A Generic Framework for LLM Safety Evaluation
di: Jindal, Madhur, et al.
Pubblicazione: (2025)
di: Jindal, Madhur, et al.
Pubblicazione: (2025)
Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking Attacks
di: Saha, Shoumik, et al.
Pubblicazione: (2025)
di: Saha, Shoumik, et al.
Pubblicazione: (2025)
Information Security Based on LLM Approaches: A Review
di: Gong, Chang, et al.
Pubblicazione: (2025)
di: Gong, Chang, et al.
Pubblicazione: (2025)
Security awareness in LLM agents: the NDAI zone case
di: Bottazzi, Enrico, et al.
Pubblicazione: (2026)
di: Bottazzi, Enrico, et al.
Pubblicazione: (2026)
The Authorization-Execution Gap Is a Major Safety and Security Problem in Open-World Agents
di: Wu, Baoyuan, et al.
Pubblicazione: (2026)
di: Wu, Baoyuan, et al.
Pubblicazione: (2026)
Documenti analoghi
-
DeBackdoor: A Deductive Framework for Detecting Backdoor Attacks on Deep Models with Limited Data
di: Popovic, Dorde, et al.
Pubblicazione: (2025) -
CallShield: Secure Caller Authentication over Real-Time Audio Channels
di: Rabh, Mouna, et al.
Pubblicazione: (2026) -
Swiss-Bench 003: Evaluating LLM Reliability and Adversarial Security for Swiss Regulatory Contexts
di: Uenal, Fatih
Pubblicazione: (2026) -
MANTIS: Detection of Zero-Day Malicious Domains Leveraging Low Reputed Hosting Infrastructure
di: Deniz, Fatih, et al.
Pubblicazione: (2025) -
Simplified and Secure MCP Gateways for Enterprise AI Integration
di: Brett, Ivo
Pubblicazione: (2025)