Generative AI Security: Challenges and Countermeasures
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhu, Banghua, Mu, Norman, Jiao, Jiantao, Wagner, David |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PAL: Proxy-Guided Black-Box Attack on Large Language Models
von: Sitawarin, Chawin, et al.
Veröffentlicht: (2024)
von: Sitawarin, Chawin, et al.
Veröffentlicht: (2024)
LLM Platform Security: Applying a Systematic Evaluation Framework to OpenAI's ChatGPT Plugins
von: Iqbal, Umar, et al.
Veröffentlicht: (2023)
von: Iqbal, Umar, et al.
Veröffentlicht: (2023)
A Survey of Privacy-Preserving Model Explanations: Privacy Risks, Attacks, and Countermeasures
von: Nguyen, Thanh Tam, et al.
Veröffentlicht: (2024)
von: Nguyen, Thanh Tam, et al.
Veröffentlicht: (2024)
Urania: Differentially Private Insights into AI Use
von: Liu, Daogao, et al.
Veröffentlicht: (2025)
von: Liu, Daogao, et al.
Veröffentlicht: (2025)
Clio: Privacy-Preserving Insights into Real-World AI Use
von: Tamkin, Alex, et al.
Veröffentlicht: (2024)
von: Tamkin, Alex, et al.
Veröffentlicht: (2024)
SecGenAI: Enhancing Security of Cloud-based Generative AI Applications within Australian Critical Technologies of National Interest
von: Haryanto, Christoforus Yoga, et al.
Veröffentlicht: (2024)
von: Haryanto, Christoforus Yoga, et al.
Veröffentlicht: (2024)
Can a large language model be a gaslighter?
von: Li, Wei, et al.
Veröffentlicht: (2024)
von: Li, Wei, et al.
Veröffentlicht: (2024)
An In-Depth Investigation of Data Collection in LLM App Ecosystems
von: Wu, Yuhao, et al.
Veröffentlicht: (2024)
von: Wu, Yuhao, et al.
Veröffentlicht: (2024)
Superficial Safety Alignment Hypothesis
von: Li, Jianwei, et al.
Veröffentlicht: (2024)
von: Li, Jianwei, et al.
Veröffentlicht: (2024)
IsolateGPT: An Execution Isolation Architecture for LLM-Based Agentic Systems
von: Wu, Yuhao, et al.
Veröffentlicht: (2024)
von: Wu, Yuhao, et al.
Veröffentlicht: (2024)
Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
von: Zhang, Andy K., et al.
Veröffentlicht: (2024)
von: Zhang, Andy K., et al.
Veröffentlicht: (2024)
Misaligned Roles, Misplaced Images: Structural Input Perturbations Expose Multimodal Alignment Blind Spots
von: Shayegani, Erfan, et al.
Veröffentlicht: (2025)
von: Shayegani, Erfan, et al.
Veröffentlicht: (2025)
Just Do It!? Computer-Use Agents Exhibit Blind Goal-Directedness
von: Shayegani, Erfan, et al.
Veröffentlicht: (2025)
von: Shayegani, Erfan, et al.
Veröffentlicht: (2025)
What Makes an Evaluation Useful? Common Pitfalls and Best Practices
von: Gekker, Gil, et al.
Veröffentlicht: (2025)
von: Gekker, Gil, et al.
Veröffentlicht: (2025)
Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods
von: Jang, Yeonwoo, et al.
Veröffentlicht: (2025)
von: Jang, Yeonwoo, et al.
Veröffentlicht: (2025)
LLM Security: Vulnerabilities, Attacks, Defenses, and Countermeasures
von: Aguilera-Martínez, Francisco, et al.
Veröffentlicht: (2025)
von: Aguilera-Martínez, Francisco, et al.
Veröffentlicht: (2025)
SoK: On the Offensive Potential of AI
von: Schröer, Saskia Laura, et al.
Veröffentlicht: (2024)
von: Schröer, Saskia Laura, et al.
Veröffentlicht: (2024)
SecureCode: A Production-Grade Multi-Turn Dataset for Training Security-Aware Code Generation Models
von: Thornton, Scott
Veröffentlicht: (2025)
von: Thornton, Scott
Veröffentlicht: (2025)
Privacy at a Price: Exploring its Dual Impact on AI Fairness
von: Yang, Mengmeng, et al.
Veröffentlicht: (2024)
von: Yang, Mengmeng, et al.
Veröffentlicht: (2024)
Mark My Words: Analyzing and Evaluating Language Model Watermarks
von: Piet, Julien, et al.
Veröffentlicht: (2023)
von: Piet, Julien, et al.
Veröffentlicht: (2023)
Towards Optimal Statistical Watermarking
von: Huang, Baihe, et al.
Veröffentlicht: (2023)
von: Huang, Baihe, et al.
Veröffentlicht: (2023)
Secure Multi-Modal Data Fusion in Federated Digital Health Systems via MCP
von: Aueawatthanaphisut, Aueaphum
Veröffentlicht: (2025)
von: Aueawatthanaphisut, Aueaphum
Veröffentlicht: (2025)
Challenges in Ensuring AI Safety in DeepSeek-R1 Models: The Shortcomings of Reinforcement Learning Strategies
von: Parmar, Manojkumar, et al.
Veröffentlicht: (2025)
von: Parmar, Manojkumar, et al.
Veröffentlicht: (2025)
NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security
von: Shao, Minghao, et al.
Veröffentlicht: (2024)
von: Shao, Minghao, et al.
Veröffentlicht: (2024)
Gandalf the Red: Adaptive Security for LLMs
von: Pfister, Niklas, et al.
Veröffentlicht: (2025)
von: Pfister, Niklas, et al.
Veröffentlicht: (2025)
Security Degradation in Iterative AI Code Generation -- A Systematic Analysis of the Paradox
von: Shukla, Shivani, et al.
Veröffentlicht: (2025)
von: Shukla, Shivani, et al.
Veröffentlicht: (2025)
Making AI-Assisted Grant Evaluation Auditable without Exposing the Model
von: Bicakci, Kemal
Veröffentlicht: (2026)
von: Bicakci, Kemal
Veröffentlicht: (2026)
SecEncoder: Logs are All You Need in Security
von: Bulut, Muhammed Fatih, et al.
Veröffentlicht: (2024)
von: Bulut, Muhammed Fatih, et al.
Veröffentlicht: (2024)
IH-Challenge: A Training Dataset to Improve Instruction Hierarchy on Frontier LLMs
von: Guo, Chuan, et al.
Veröffentlicht: (2026)
von: Guo, Chuan, et al.
Veröffentlicht: (2026)
ROK-FORTRESS: Measuring the Effect of Geopolitical Transcreation for National Security and Public Safety
von: Lee, Michael S., et al.
Veröffentlicht: (2026)
von: Lee, Michael S., et al.
Veröffentlicht: (2026)
SecureBreak -- A dataset towards safe and secure models
von: Arazzi, Marco, et al.
Veröffentlicht: (2026)
von: Arazzi, Marco, et al.
Veröffentlicht: (2026)
Children's Voice Privacy: First Steps And Emerging Challenges
von: Kulkarni, Ajinkya, et al.
Veröffentlicht: (2025)
von: Kulkarni, Ajinkya, et al.
Veröffentlicht: (2025)
Enhancing Prompt Injection Attacks to LLMs via Poisoning Alignment
von: Shao, Zedian, et al.
Veröffentlicht: (2024)
von: Shao, Zedian, et al.
Veröffentlicht: (2024)
Benchmark Early and Red Team Often: A Framework for Assessing and Managing Dual-Use Hazards of AI Foundation Models
von: Barrett, Anthony M., et al.
Veröffentlicht: (2024)
von: Barrett, Anthony M., et al.
Veröffentlicht: (2024)
How Well Can LLM Agents Simulate End-User Security and Privacy Attitudes and Behaviors?
von: Li, Yuxuan, et al.
Veröffentlicht: (2026)
von: Li, Yuxuan, et al.
Veröffentlicht: (2026)
Securing Multi-turn Conversational Language Models From Distributed Backdoor Triggers
von: Tong, Terry, et al.
Veröffentlicht: (2024)
von: Tong, Terry, et al.
Veröffentlicht: (2024)
Covert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation
von: Halawi, Danny, et al.
Veröffentlicht: (2024)
von: Halawi, Danny, et al.
Veröffentlicht: (2024)
Taylor Unswift: Secured Weight Release for Large Language Models via Taylor Expansion
von: Wang, Guanchu, et al.
Veröffentlicht: (2024)
von: Wang, Guanchu, et al.
Veröffentlicht: (2024)
Noise Contrastive Estimation-based Matching Framework for Low-Resource Security Attack Pattern Recognition
von: Nguyen, Tu, et al.
Veröffentlicht: (2024)
von: Nguyen, Tu, et al.
Veröffentlicht: (2024)
Adversarial Intent is a Latent Variable: Stateful Trust Inference for Securing Multimodal Agentic RAG
von: Singh, Inderjeet, et al.
Veröffentlicht: (2026)
von: Singh, Inderjeet, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
PAL: Proxy-Guided Black-Box Attack on Large Language Models
von: Sitawarin, Chawin, et al.
Veröffentlicht: (2024) -
LLM Platform Security: Applying a Systematic Evaluation Framework to OpenAI's ChatGPT Plugins
von: Iqbal, Umar, et al.
Veröffentlicht: (2023) -
A Survey of Privacy-Preserving Model Explanations: Privacy Risks, Attacks, and Countermeasures
von: Nguyen, Thanh Tam, et al.
Veröffentlicht: (2024) -
Urania: Differentially Private Insights into AI Use
von: Liu, Daogao, et al.
Veröffentlicht: (2025) -
Clio: Privacy-Preserving Insights into Real-World AI Use
von: Tamkin, Alex, et al.
Veröffentlicht: (2024)