AI Security in the Foundation Model Era: A Comprehensive Survey from a Unified Perspective
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Zhenyi, Luan, Siyu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AgentDoG 1.5: A Lightweight and Scalable Alignment Framework for AI Agent Safety and Security
by: Liu, Dongrui, et al.
Published: (2026)
by: Liu, Dongrui, et al.
Published: (2026)
On the Trustworthiness Landscape of State-of-the-art Generative Models: A Survey and Outlook
by: Fan, Mingyuan, et al.
Published: (2023)
by: Fan, Mingyuan, et al.
Published: (2023)
Beyond Boundaries: A Comprehensive Survey of Transferable Attacks on AI Systems
by: Wang, Guangjing, et al.
Published: (2023)
by: Wang, Guangjing, et al.
Published: (2023)
OmniGuard: Unified Omni-Modal Guardrails with Deliberate Reasoning
by: Zhu, Boyu, et al.
Published: (2025)
by: Zhu, Boyu, et al.
Published: (2025)
Watermark-based Attribution of AI-Generated Content
by: Jiang, Zhengyuan, et al.
Published: (2024)
by: Jiang, Zhengyuan, et al.
Published: (2024)
Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety
by: Ma, Xingjun, et al.
Published: (2025)
by: Ma, Xingjun, et al.
Published: (2025)
MOSSBench: Is Your Multimodal Language Model Oversensitive to Safe Queries?
by: Li, Xirui, et al.
Published: (2024)
by: Li, Xirui, et al.
Published: (2024)
Do We Really Need Curated Malicious Data for Safety Alignment in Multi-modal Large Language Models?
by: Wang, Yanbo, et al.
Published: (2025)
by: Wang, Yanbo, et al.
Published: (2025)
Membership Inference Attacks against Large Vision-Language Models
by: Li, Zhan, et al.
Published: (2024)
by: Li, Zhan, et al.
Published: (2024)
The Curse of Recursion: Training on Generated Data Makes Models Forget
by: Shumailov, Ilia, et al.
Published: (2023)
by: Shumailov, Ilia, et al.
Published: (2023)
T2VSafetyBench: Evaluating the Safety of Text-to-Video Generative Models
by: Miao, Yibo, et al.
Published: (2024)
by: Miao, Yibo, et al.
Published: (2024)
Failures to Find Transferable Image Jailbreaks Between Vision-Language Models
by: Schaeffer, Rylan, et al.
Published: (2024)
by: Schaeffer, Rylan, et al.
Published: (2024)
The Safety Reminder: A Soft Prompt to Reactivate Delayed Safety Awareness in Vision-Language Models
by: Tang, Peiyuan, et al.
Published: (2025)
by: Tang, Peiyuan, et al.
Published: (2025)
To See is Not to Learn: Protecting Multimodal Data from Unauthorized Fine-Tuning of Large Vision-Language Model
by: Zhao, Chengshuai, et al.
Published: (2026)
by: Zhao, Chengshuai, et al.
Published: (2026)
Task-Agnostic Attacks Against Vision Foundation Models
by: Pulfer, Brian, et al.
Published: (2025)
by: Pulfer, Brian, et al.
Published: (2025)
The Salami Slicing Threat: Exploiting Cumulative Risks in LLM Systems
by: Zhang, Yihao, et al.
Published: (2026)
by: Zhang, Yihao, et al.
Published: (2026)
SecureScan: An AI-Driven Multi-Layer Framework for Malware and Phishing Detection Using Logistic Regression and Threat Intelligence Integration
by: Firdos, Rumman, et al.
Published: (2026)
by: Firdos, Rumman, et al.
Published: (2026)
Towards Personalized Federated Learning via Comprehensive Knowledge Distillation
by: Wang, Pengju, et al.
Published: (2024)
by: Wang, Pengju, et al.
Published: (2024)
When Think-with-Image Meets Safety: What Determines Multimodal Jailbreak Robustness?
by: Tian, Yuan, et al.
Published: (2026)
by: Tian, Yuan, et al.
Published: (2026)
Solving Trojan Detection Competitions with Linear Weight Classification
by: Huster, Todd, et al.
Published: (2024)
by: Huster, Todd, et al.
Published: (2024)
X-Boundary: Establishing Exact Safety Boundary to Shield LLMs from Multi-Turn Jailbreaks without Compromising Usability
by: Lu, Xiaoya, et al.
Published: (2025)
by: Lu, Xiaoya, et al.
Published: (2025)
ArtPerception: ASCII Art-based Jailbreak on LLMs with Recognition Pre-test
by: Yang, Guan-Yan, et al.
Published: (2025)
by: Yang, Guan-Yan, et al.
Published: (2025)
Differentially Private Bias-Term Fine-tuning of Foundation Models
by: Bu, Zhiqi, et al.
Published: (2022)
by: Bu, Zhiqi, et al.
Published: (2022)
A Survey on Physical Adversarial Attacks against Face Recognition Systems
by: Wang, Mingsi, et al.
Published: (2024)
by: Wang, Mingsi, et al.
Published: (2024)
Scalable Secure Biometric Authentication without Auxiliary Identifiers
by: Bienstock, Alexander, et al.
Published: (2026)
by: Bienstock, Alexander, et al.
Published: (2026)
Rethinking Data Protection in the (Generative) Artificial Intelligence Era
by: Li, Yiming, et al.
Published: (2025)
by: Li, Yiming, et al.
Published: (2025)
An Efficient and Multi-private Key Secure Aggregation for Federated Learning
by: Yang, Xue, et al.
Published: (2023)
by: Yang, Xue, et al.
Published: (2023)
AdvSecureNet: A Python Toolkit for Adversarial Machine Learning
by: Catal, Melih, et al.
Published: (2024)
by: Catal, Melih, et al.
Published: (2024)
Securing Visually-Aware Recommender Systems: An Adversarial Image Reconstruction and Detection Framework
by: Yin, Minglei, et al.
Published: (2023)
by: Yin, Minglei, et al.
Published: (2023)
Inference Attacks: A Taxonomy, Survey, and Promising Directions
by: Wu, Feng, et al.
Published: (2024)
by: Wu, Feng, et al.
Published: (2024)
From Attack to Defense: Insights into Deep Learning Security Measures in Black-Box Settings
by: Juraev, Firuz, et al.
Published: (2024)
by: Juraev, Firuz, et al.
Published: (2024)
A Survey on the Safety and Security Threats of Computer-Using Agents: JARVIS or Ultron?
by: Chen, Ada, et al.
Published: (2025)
by: Chen, Ada, et al.
Published: (2025)
On the explainable properties of 1-Lipschitz Neural Networks: An Optimal Transport Perspective
by: Serrurier, Mathieu, et al.
Published: (2022)
by: Serrurier, Mathieu, et al.
Published: (2022)
JailbreakZoo: Survey, Landscapes, and Horizons in Jailbreaking Large Language and Vision-Language Models
by: Jin, Haibo, et al.
Published: (2024)
by: Jin, Haibo, et al.
Published: (2024)
A Survey on the Application of Generative Adversarial Networks in Cybersecurity: Prospective, Direction and Open Research Scopes
by: Arifin, Md Mashrur, et al.
Published: (2024)
by: Arifin, Md Mashrur, et al.
Published: (2024)
EditTrack: Detecting and Attributing AI-assisted Image Editing
by: Jiang, Zhengyuan, et al.
Published: (2025)
by: Jiang, Zhengyuan, et al.
Published: (2025)
Crafting Adversarial Inputs for Large Vision-Language Models Using Black-Box Optimization
by: Guan, Jiwei, et al.
Published: (2026)
by: Guan, Jiwei, et al.
Published: (2026)
Beyond Binary Classification: A Semi-supervised Approach to Generalized AI-generated Image Detection
by: Nguyen-Le, Hong-Hanh, et al.
Published: (2025)
by: Nguyen-Le, Hong-Hanh, et al.
Published: (2025)
FreezeAsGuard: Mitigating Illegal Adaptation of Diffusion Models via Selective Tensor Freezing
by: Huang, Kai, et al.
Published: (2024)
by: Huang, Kai, et al.
Published: (2024)
Towards Resilient Safety-driven Unlearning for Diffusion Models against Downstream Fine-tuning
by: Li, Boheng, et al.
Published: (2025)
by: Li, Boheng, et al.
Published: (2025)
Similar Items
-
AgentDoG 1.5: A Lightweight and Scalable Alignment Framework for AI Agent Safety and Security
by: Liu, Dongrui, et al.
Published: (2026) -
On the Trustworthiness Landscape of State-of-the-art Generative Models: A Survey and Outlook
by: Fan, Mingyuan, et al.
Published: (2023) -
Beyond Boundaries: A Comprehensive Survey of Transferable Attacks on AI Systems
by: Wang, Guangjing, et al.
Published: (2023) -
OmniGuard: Unified Omni-Modal Guardrails with Deliberate Reasoning
by: Zhu, Boyu, et al.
Published: (2025) -
Watermark-based Attribution of AI-Generated Content
by: Jiang, Zhengyuan, et al.
Published: (2024)