SecureBreak -- A dataset towards safe and secure models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Arazzi, Marco, Kembu, Vignesh Kumar, Nocera, Antonino |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
XBreaking: Understanding how LLMs security alignment can be broken
von: Arazzi, Marco, et al.
Veröffentlicht: (2025)
von: Arazzi, Marco, et al.
Veröffentlicht: (2025)
Privacy Preserving and Robust Aggregation for Cross-Silo Federated Learning in Non-IID Settings
von: Arazzi, Marco, et al.
Veröffentlicht: (2025)
von: Arazzi, Marco, et al.
Veröffentlicht: (2025)
Security in LLM-as-a-Judge: A Comprehensive SoK
von: Masoud, Aiman Al, et al.
Veröffentlicht: (2026)
von: Masoud, Aiman Al, et al.
Veröffentlicht: (2026)
A Deep Reinforcement Learning Approach for Security-Aware Service Acquisition in IoT
von: Arazzi, Marco, et al.
Veröffentlicht: (2024)
von: Arazzi, Marco, et al.
Veröffentlicht: (2024)
LoRA as Oracle
von: Arazzi, Marco, et al.
Veröffentlicht: (2026)
von: Arazzi, Marco, et al.
Veröffentlicht: (2026)
A Novel IoT Trust Model Leveraging Fully Distributed Behavioral Fingerprinting and Secure Delegation
von: Arazzi, Marco, et al.
Veröffentlicht: (2023)
von: Arazzi, Marco, et al.
Veröffentlicht: (2023)
Secure Federated Data Distillation
von: Arazzi, Marco, et al.
Veröffentlicht: (2025)
von: Arazzi, Marco, et al.
Veröffentlicht: (2025)
You Snooze, You Lose: Automatic Safety Alignment Restoration through Neural Weight Translation
von: Arazzi, Marco, et al.
Veröffentlicht: (2026)
von: Arazzi, Marco, et al.
Veröffentlicht: (2026)
KDk: A Defense Mechanism Against Label Inference Attacks in Vertical Federated Learning
von: Arazzi, Marco, et al.
Veröffentlicht: (2024)
von: Arazzi, Marco, et al.
Veröffentlicht: (2024)
SD-RAG: A Prompt-Injection-Resilient Framework for Selective Disclosure in Retrieval-Augmented Generation
von: Masoud, Aiman Al, et al.
Veröffentlicht: (2026)
von: Masoud, Aiman Al, et al.
Veröffentlicht: (2026)
Subject Data Auditing via Source Inference Attack in Cross-Silo Federated Learning
von: Li, Jiaxin, et al.
Veröffentlicht: (2024)
von: Li, Jiaxin, et al.
Veröffentlicht: (2024)
Privacy-Preserving in Blockchain-based Federated Learning Systems
von: M., Sameera K., et al.
Veröffentlicht: (2024)
von: M., Sameera K., et al.
Veröffentlicht: (2024)
Let's Focus: Focused Backdoor Attack against Federated Transfer Learning
von: Arazzi, Marco, et al.
Veröffentlicht: (2024)
von: Arazzi, Marco, et al.
Veröffentlicht: (2024)
The Ethics of Interaction: Mitigating Security Threats in LLMs
von: Kumar, Ashutosh, et al.
Veröffentlicht: (2024)
von: Kumar, Ashutosh, et al.
Veröffentlicht: (2024)
Gandalf the Red: Adaptive Security for LLMs
von: Pfister, Niklas, et al.
Veröffentlicht: (2025)
von: Pfister, Niklas, et al.
Veröffentlicht: (2025)
SecureCode: A Production-Grade Multi-Turn Dataset for Training Security-Aware Code Generation Models
von: Thornton, Scott
Veröffentlicht: (2025)
von: Thornton, Scott
Veröffentlicht: (2025)
What is in Your Safe Data? Identifying Benign Data that Breaks Safety
von: He, Luxi, et al.
Veröffentlicht: (2024)
von: He, Luxi, et al.
Veröffentlicht: (2024)
Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM
von: Cao, Bochuan, et al.
Veröffentlicht: (2023)
von: Cao, Bochuan, et al.
Veröffentlicht: (2023)
Toxicity Detection towards Adaptability to Changing Perturbations
von: Kang, Hankun, et al.
Veröffentlicht: (2024)
von: Kang, Hankun, et al.
Veröffentlicht: (2024)
SequentialBreak: Large Language Models Can be Fooled by Embedding Jailbreak Prompts into Sequential Prompt Chains
von: Saiem, Bijoy Ahmed, et al.
Veröffentlicht: (2024)
von: Saiem, Bijoy Ahmed, et al.
Veröffentlicht: (2024)
SALT: Steering Activations towards Leakage-free Thinking in Chain of Thought
von: Batra, Shourya, et al.
Veröffentlicht: (2025)
von: Batra, Shourya, et al.
Veröffentlicht: (2025)
SecEncoder: Logs are All You Need in Security
von: Bulut, Muhammed Fatih, et al.
Veröffentlicht: (2024)
von: Bulut, Muhammed Fatih, et al.
Veröffentlicht: (2024)
PII-Compass: Guiding LLM training data extraction prompts towards the target PII via grounding
von: Nakka, Krishna Kanth, et al.
Veröffentlicht: (2024)
von: Nakka, Krishna Kanth, et al.
Veröffentlicht: (2024)
Securing Multi-turn Conversational Language Models From Distributed Backdoor Triggers
von: Tong, Terry, et al.
Veröffentlicht: (2024)
von: Tong, Terry, et al.
Veröffentlicht: (2024)
Taylor Unswift: Secured Weight Release for Large Language Models via Taylor Expansion
von: Wang, Guanchu, et al.
Veröffentlicht: (2024)
von: Wang, Guanchu, et al.
Veröffentlicht: (2024)
When Forgetting Triggers Backdoors: A Clean Unlearning Attack
von: Arazzi, Marco, et al.
Veröffentlicht: (2025)
von: Arazzi, Marco, et al.
Veröffentlicht: (2025)
Adversarial Intent is a Latent Variable: Stateful Trust Inference for Securing Multimodal Agentic RAG
von: Singh, Inderjeet, et al.
Veröffentlicht: (2026)
von: Singh, Inderjeet, et al.
Veröffentlicht: (2026)
SentinelLMs: Encrypted Input Adaptation and Fine-tuning of Language Models for Private and Secure Inference
von: Mishra, Abhijit, et al.
Veröffentlicht: (2023)
von: Mishra, Abhijit, et al.
Veröffentlicht: (2023)
Noise Contrastive Estimation-based Matching Framework for Low-Resource Security Attack Pattern Recognition
von: Nguyen, Tu, et al.
Veröffentlicht: (2024)
von: Nguyen, Tu, et al.
Veröffentlicht: (2024)
Label Inference Attacks against Node-level Vertical Federated GNNs
von: Arazzi, Marco, et al.
Veröffentlicht: (2023)
von: Arazzi, Marco, et al.
Veröffentlicht: (2023)
Less Data, More Security: Advancing Cybersecurity LLMs Specialization via Resource-Efficient Domain-Adaptive Continuous Pre-training with Minimal Tokens
von: Salahuddin, Salahuddin, et al.
Veröffentlicht: (2025)
von: Salahuddin, Salahuddin, et al.
Veröffentlicht: (2025)
Attack and defense techniques in large language models: A survey and new perspectives
von: Liao, Zhiyu, et al.
Veröffentlicht: (2025)
von: Liao, Zhiyu, et al.
Veröffentlicht: (2025)
Generative AI Security: Challenges and Countermeasures
von: Zhu, Banghua, et al.
Veröffentlicht: (2024)
von: Zhu, Banghua, et al.
Veröffentlicht: (2024)
PromptScreen: Efficient Jailbreak Mitigation Using Semantic Linear Classification in a Multi-Staged Pipeline
von: Rao, Akshaj Prashanth, et al.
Veröffentlicht: (2025)
von: Rao, Akshaj Prashanth, et al.
Veröffentlicht: (2025)
Jailbreaking Large Language Models with Symbolic Mathematics
von: Bethany, Emet, et al.
Veröffentlicht: (2024)
von: Bethany, Emet, et al.
Veröffentlicht: (2024)
Certifying LLM Safety against Adversarial Prompting
von: Kumar, Aounon, et al.
Veröffentlicht: (2023)
von: Kumar, Aounon, et al.
Veröffentlicht: (2023)
EPSVec: Efficient and Private Synthetic Data Generation via Dataset Vectors
von: Banayeeanzade, Amin, et al.
Veröffentlicht: (2026)
von: Banayeeanzade, Amin, et al.
Veröffentlicht: (2026)
LLM Platform Security: Applying a Systematic Evaluation Framework to OpenAI's ChatGPT Plugins
von: Iqbal, Umar, et al.
Veröffentlicht: (2023)
von: Iqbal, Umar, et al.
Veröffentlicht: (2023)
TableGuard -- Securing Structured & Unstructured Data
von: Sharma, Anantha, et al.
Veröffentlicht: (2024)
von: Sharma, Anantha, et al.
Veröffentlicht: (2024)
Security Degradation in Iterative AI Code Generation -- A Systematic Analysis of the Paradox
von: Shukla, Shivani, et al.
Veröffentlicht: (2025)
von: Shukla, Shivani, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
XBreaking: Understanding how LLMs security alignment can be broken
von: Arazzi, Marco, et al.
Veröffentlicht: (2025) -
Privacy Preserving and Robust Aggregation for Cross-Silo Federated Learning in Non-IID Settings
von: Arazzi, Marco, et al.
Veröffentlicht: (2025) -
Security in LLM-as-a-Judge: A Comprehensive SoK
von: Masoud, Aiman Al, et al.
Veröffentlicht: (2026) -
A Deep Reinforcement Learning Approach for Security-Aware Service Acquisition in IoT
von: Arazzi, Marco, et al.
Veröffentlicht: (2024) -
LoRA as Oracle
von: Arazzi, Marco, et al.
Veröffentlicht: (2026)