ARMOR: Aligning Secure and Safe Large Language Models via Meticulous Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Zhengyue, Ma, Yingzi, Jha, Somesh, Pavone, Marco, McDaniel, Patrick, Xiao, Chaowei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A New Era in LLM Security: Exploring Security Concerns in Real-World LLM-based Systems
by: Wu, Fangzhou, et al.
Published: (2024)
by: Wu, Fangzhou, et al.
Published: (2024)
PRISM: Robust VLM Alignment with Principled Reasoning for Integrated Safety in Multimodality
by: Li, Nanxi, et al.
Published: (2025)
by: Li, Nanxi, et al.
Published: (2025)
LPG: Balancing Efficiency and Policy Reasoning in Latent Policy Guardrails
by: Li, Nanxi, et al.
Published: (2026)
by: Li, Nanxi, et al.
Published: (2026)
Doxing via the Lens: Revealing Location-related Privacy Leakage on Multi-modal Large Reasoning Models
by: Luo, Weidi, et al.
Published: (2025)
by: Luo, Weidi, et al.
Published: (2025)
AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs
by: Liu, Xiaogeng, et al.
Published: (2024)
by: Liu, Xiaogeng, et al.
Published: (2024)
A Public and Reproducible Assessment of the Topics API on Real Data
by: Beugin, Yohan, et al.
Published: (2024)
by: Beugin, Yohan, et al.
Published: (2024)
Technical Report: The Need for a (Research) Sandstorm through the Privacy Sandbox
by: Beugin, Yohan, et al.
Published: (2025)
by: Beugin, Yohan, et al.
Published: (2025)
Confused ChatGPT: Cross-App Context Poisoning via First-Party APIs
by: Wang, Chao, et al.
Published: (2026)
by: Wang, Chao, et al.
Published: (2026)
It's a Feature, Not a Bug: Secure and Auditable State Rollback for Confidential Cloud Applications
by: Burke, Quinn, et al.
Published: (2025)
by: Burke, Quinn, et al.
Published: (2025)
On Scalable Integrity Checking for Secure Cloud Disks
by: Burke, Quinn, et al.
Published: (2024)
by: Burke, Quinn, et al.
Published: (2024)
Targeting Alignment: Extracting Safety Classifiers of Aligned LLMs
by: Ferrand, Jean-Charles Noirot, et al.
Published: (2025)
by: Ferrand, Jean-Charles Noirot, et al.
Published: (2025)
Secure IP Address Allocation at Cloud Scale
by: Pauley, Eric, et al.
Published: (2022)
by: Pauley, Eric, et al.
Published: (2022)
The Role of Learning in Attacking ML-based Network Intrusion Detection
by: Domico, Kyle, et al.
Published: (2026)
by: Domico, Kyle, et al.
Published: (2026)
Dependency-Aware Privacy for Multi-turn Agents
by: Anshumaan, Divyam, et al.
Published: (2026)
by: Anshumaan, Divyam, et al.
Published: (2026)
SLVR: Securely Leveraging Client Validation for Robust Federated Learning
by: Choi, Jihye, et al.
Published: (2025)
by: Choi, Jihye, et al.
Published: (2025)
AutoDAN-Reasoning: Enhancing Strategies Exploration based Jailbreak Attacks with Test-Time Scaling
by: Liu, Xiaogeng, et al.
Published: (2025)
by: Liu, Xiaogeng, et al.
Published: (2025)
Longitudinal Analyses of SAST Tools: A CodeQL Case Study
by: Ferrand, Jean-Charles Noirot, et al.
Published: (2026)
by: Ferrand, Jean-Charles Noirot, et al.
Published: (2026)
Deserialization Gadget Chains are not a Pathological Problem in Android:an In-Depth Study of Java Gadget Chains in AOSP
by: Kreyssig, Bruno, et al.
Published: (2025)
by: Kreyssig, Bruno, et al.
Published: (2025)
Characterizing the Modification Space of Signature IDS Rules
by: Guide, Ryan, et al.
Published: (2024)
by: Guide, Ryan, et al.
Published: (2024)
PRP: Propagating Universal Perturbations to Attack Large Language Model Guard-Rails
by: Mangaokar, Neal, et al.
Published: (2024)
by: Mangaokar, Neal, et al.
Published: (2024)
Mitigating Fine-tuning based Jailbreak Attack with Backdoor Enhanced Safety Alignment
by: Wang, Jiongxiao, et al.
Published: (2024)
by: Wang, Jiongxiao, et al.
Published: (2024)
SafeClaw-R: Towards Safe and Secure Multi-Agent Personal Assistants
by: Wang, Haoyu, et al.
Published: (2026)
by: Wang, Haoyu, et al.
Published: (2026)
ReasAlign: Reasoning Enhanced Safety Alignment against Prompt Injection Attack
by: Li, Hao, et al.
Published: (2026)
by: Li, Hao, et al.
Published: (2026)
LibIHT: A Hardware-Based Approach to Efficient and Evasion-Resistant Dynamic Binary Analysis
by: Zhao, Changyu, et al.
Published: (2025)
by: Zhao, Changyu, et al.
Published: (2025)
Err on the Side of Texture: Texture Bias on Real Data
by: Hoak, Blaine, et al.
Published: (2024)
by: Hoak, Blaine, et al.
Published: (2024)
AdaShield: Safeguarding Multimodal Large Language Models from Structure-based Attack via Adaptive Shield Prompting
by: Wang, Yu, et al.
Published: (2024)
by: Wang, Yu, et al.
Published: (2024)
AgentDyn: Are Your Agent Security Defenses Deployable in Real-World Dynamic Environments?
by: Li, Hao, et al.
Published: (2026)
by: Li, Hao, et al.
Published: (2026)
Safety Layers in Aligned Large Language Models: The Key to LLM Security
by: Li, Shen, et al.
Published: (2024)
by: Li, Shen, et al.
Published: (2024)
Systems Security Foundations for Agentic Computing
by: Christodorescu, Mihai, et al.
Published: (2025)
by: Christodorescu, Mihai, et al.
Published: (2025)
AgentSys: Secure and Dynamic LLM Agents Through Explicit Hierarchical Memory Management
by: Wen, Ruoyao, et al.
Published: (2026)
by: Wen, Ruoyao, et al.
Published: (2026)
Cooking Up Risks: Benchmarking and Reducing Food Safety Risks in Large Language Models
by: Luo, Weidi, et al.
Published: (2026)
by: Luo, Weidi, et al.
Published: (2026)
Efficient Storage Integrity in Adversarial Settings
by: Burke, Quinn, et al.
Published: (2025)
by: Burke, Quinn, et al.
Published: (2025)
PSRT: Accelerating LRM-based Guard Models via Prefilled Safe Reasoning Traces
by: Zhao, Jiawei, et al.
Published: (2025)
by: Zhao, Jiawei, et al.
Published: (2025)
Securing Cloud File Systems with Trusted Execution
by: Burke, Quinn, et al.
Published: (2023)
by: Burke, Quinn, et al.
Published: (2023)
A Practical Guideline and Taxonomy to LLVM's Control Flow Integrity
by: Houy, Sabine, et al.
Published: (2025)
by: Houy, Sabine, et al.
Published: (2025)
ReasoningBomb: A Stealthy Denial-of-Service Attack by Inducing Pathologically Long Reasoning in Large Reasoning Models
by: Liu, Xiaogeng, et al.
Published: (2026)
by: Liu, Xiaogeng, et al.
Published: (2026)
System-Level Defense against Indirect Prompt Injection Attacks: An Information Flow Control Perspective
by: Wu, Fangzhou, et al.
Published: (2024)
by: Wu, Fangzhou, et al.
Published: (2024)
An Approach for Safe and Secure Software Protection Supported by Symbolic Execution
by: Dorfmeister, Daniel, et al.
Published: (2026)
by: Dorfmeister, Daniel, et al.
Published: (2026)
ARMOR: Shielding Unlearnable Examples against Data Augmentation
by: Gong, Xueluan, et al.
Published: (2025)
by: Gong, Xueluan, et al.
Published: (2025)
PolicyLR: A Logic Representation For Privacy Policies
by: Hooda, Ashish, et al.
Published: (2024)
by: Hooda, Ashish, et al.
Published: (2024)
Similar Items
-
A New Era in LLM Security: Exploring Security Concerns in Real-World LLM-based Systems
by: Wu, Fangzhou, et al.
Published: (2024) -
PRISM: Robust VLM Alignment with Principled Reasoning for Integrated Safety in Multimodality
by: Li, Nanxi, et al.
Published: (2025) -
LPG: Balancing Efficiency and Policy Reasoning in Latent Policy Guardrails
by: Li, Nanxi, et al.
Published: (2026) -
Doxing via the Lens: Revealing Location-related Privacy Leakage on Multi-modal Large Reasoning Models
by: Luo, Weidi, et al.
Published: (2025) -
AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs
by: Liu, Xiaogeng, et al.
Published: (2024)