Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework
Fuente:
arXiv
Saved in:
| Main Authors: | Dassanayake, Rishane, Demetroudi, Mario, Walpole, James, Lentati, Lindley, Brown, Jason R., Young, Edward James |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Risk Psychology & Cyber-Attack Tactics
by: Kim, Rubens, et al.
Published: (2025)
by: Kim, Rubens, et al.
Published: (2025)
Professor X: Manipulating EEG BCI with Invisible and Robust Backdoor Attack
by: Liu, Xuan-Hao, et al.
Published: (2024)
by: Liu, Xuan-Hao, et al.
Published: (2024)
Benchmarking and Understanding Safety Risks in AI Character Platforms
by: Wei, Yiluo, et al.
Published: (2025)
by: Wei, Yiluo, et al.
Published: (2025)
SoK: Come Together -- Unifying Security, Information Theory, and Cognition for a Mixed Reality Deception Attack Ontology & Analysis Framework
by: Teymourian, Ali, et al.
Published: (2025)
by: Teymourian, Ali, et al.
Published: (2025)
Multiverse Privacy Theory for Contextual Risks in Complex User-AI Interactions
by: Gumusel, Ece
Published: (2025)
by: Gumusel, Ece
Published: (2025)
Security Risks of AI Agents Hiring Humans: An Empirical Marketplace Study
by: Mehta, Pulak
Published: (2026)
by: Mehta, Pulak
Published: (2026)
Anti-Phishing Training (Still) Does Not Work: A Large-Scale Reproduction of Phishing Training Inefficacy Grounded in the NIST Phish Scale
by: Rozema, Andrew T., et al.
Published: (2025)
by: Rozema, Andrew T., et al.
Published: (2025)
Engineering Trust, Creating Vulnerability: A Socio-Technical Analysis of AI Interface Design
by: Kereopa-Yorke, Ben
Published: (2025)
by: Kereopa-Yorke, Ben
Published: (2025)
Learning from Mistakes: Can LLM Self-Recover after Misalignment?
by: Sorokoletova, Olga E., et al.
Published: (2026)
by: Sorokoletova, Olga E., et al.
Published: (2026)
Invisible, Unreadable, and Inaudible Cookie Notices: An Evaluation of Cookie Notices for Users with Visual Impairments
by: Clarke, James M., et al.
Published: (2023)
by: Clarke, James M., et al.
Published: (2023)
Analyzing Codes of Conduct for Online Safety in Video Games at Scale
by: Jiang, Jiuming, et al.
Published: (2026)
by: Jiang, Jiuming, et al.
Published: (2026)
CamLoPA: A Hidden Wireless Camera Localization Framework via Signal Propagation Path Analysis
by: Zhang, Xiang, et al.
Published: (2024)
by: Zhang, Xiang, et al.
Published: (2024)
Comparative Simulation of Phishing Attacks on a Critical Information Infrastructure Organization: An Empirical Study
by: Sirawongphatsara, Patsita, et al.
Published: (2024)
by: Sirawongphatsara, Patsita, et al.
Published: (2024)
Counterfactual Invariant Envelopes for Financial UX: Safety-Lattice Feature-Flag Governance in Crypto-Enabled Streaming
by: Malinovskiy, Anton
Published: (2026)
by: Malinovskiy, Anton
Published: (2026)
Language Model Agents Under Attack: A Cross Model-Benchmark of Profit-Seeking Behaviors in Customer Service
by: Zhang, Jingyu
Published: (2025)
by: Zhang, Jingyu
Published: (2025)
Current state of LLM Risks and AI Guardrails
by: Ayyamperumal, Suriya Ganesh, et al.
Published: (2024)
by: Ayyamperumal, Suriya Ganesh, et al.
Published: (2024)
Matcha: An IDE Plugin for Creating Accurate Privacy Nutrition Labels
by: Li, Tianshi, et al.
Published: (2024)
by: Li, Tianshi, et al.
Published: (2024)
SoK: The Privacy Paradox of Large Language Models: Advancements, Privacy Risks, and Mitigation
by: Shanmugarasa, Yashothara, et al.
Published: (2025)
by: Shanmugarasa, Yashothara, et al.
Published: (2025)
How Generative AI Empowers Attackers and Defenders Across the Trust & Safety Landscape
by: Kelley, Patrick Gage, et al.
Published: (2025)
by: Kelley, Patrick Gage, et al.
Published: (2025)
Uncovering Relationships between Android Developers, User Privacy, and Developer Willingness to Reduce Fingerprinting Risks
by: Berke, Alex, et al.
Published: (2026)
by: Berke, Alex, et al.
Published: (2026)
Understanding User Privacy Perceptions of GenAI Smartphones
by: Jin, Ran, et al.
Published: (2026)
by: Jin, Ran, et al.
Published: (2026)
CryptoGuard: An AI-Based Cryptojacking Detection Dashboard Prototype
by: Chakravorty, Amitabh, et al.
Published: (2025)
by: Chakravorty, Amitabh, et al.
Published: (2025)
Effect of Data Degradation on Motion Re-Identification
by: Nair, Vivek, et al.
Published: (2024)
by: Nair, Vivek, et al.
Published: (2024)
Actions Speak Louder Than Chats: Investigating AI Chatbot Age Gating
by: Figueira, Olivia, et al.
Published: (2026)
by: Figueira, Olivia, et al.
Published: (2026)
JailbreakLens: Visual Analysis of Jailbreak Attacks Against Large Language Models
by: Feng, Yingchaojie, et al.
Published: (2024)
by: Feng, Yingchaojie, et al.
Published: (2024)
"Tab, Tab, Bug": Security Pitfalls of Next Edit Suggestions in AI-Integrated IDEs
by: Lyu, Yunlong, et al.
Published: (2026)
by: Lyu, Yunlong, et al.
Published: (2026)
From Preventive to Reactive: How AI Coding Assistants Transform Developers' Security Awareness
by: Bappy, Faisal Haque, et al.
Published: (2026)
by: Bappy, Faisal Haque, et al.
Published: (2026)
PrivateXR: Defending Privacy Attacks in Extended Reality Through Explainable AI-Guided Differential Privacy
by: Kundu, Ripan Kumar, et al.
Published: (2025)
by: Kundu, Ripan Kumar, et al.
Published: (2025)
Cyri: A Conversational AI-based Assistant for Supporting the Human User in Detecting and Responding to Phishing Attacks
by: La Torre, Antonio, et al.
Published: (2025)
by: La Torre, Antonio, et al.
Published: (2025)
MORPHEUS: A Multidimensional Framework for Modeling, Measuring, and Mitigating Human Factors in Cybersecurity
by: Desolda, Giuseppe, et al.
Published: (2025)
by: Desolda, Giuseppe, et al.
Published: (2025)
PrivacyAssist: A User-Centric Agent Framework for Detecting Privacy Inconsistencies in Android Apps
by: Nguyen, Tran Thanh Lam, et al.
Published: (2026)
by: Nguyen, Tran Thanh Lam, et al.
Published: (2026)
From Coordinates to Context: An LLM-Bootstrapped Semantic Encoding Framework for Privacy-Preserving Mobile Sensing Stress Recognition
by: Phan, Hoang Khang, et al.
Published: (2025)
by: Phan, Hoang Khang, et al.
Published: (2025)
Effect of Duration and Delay on the Identifiability of VR Motion
by: Miller, Mark Roman, et al.
Published: (2024)
by: Miller, Mark Roman, et al.
Published: (2024)
Toward Accessible Mobile Money: A Voice-Driven, Biometrically Secured USSD Automation Framework for Visually Impaired Users
by: Ajayi, Sunday, et al.
Published: (2026)
by: Ajayi, Sunday, et al.
Published: (2026)
Defogger: A Visual Analysis Approach for Data Exploration of Sensitive Data Protected by Differential Privacy
by: Wang, Xumeng, et al.
Published: (2024)
by: Wang, Xumeng, et al.
Published: (2024)
Personalised Feedback Framework for Online Education Programmes Using Generative AI
by: Kuzminykh, Ievgeniia, et al.
Published: (2024)
by: Kuzminykh, Ievgeniia, et al.
Published: (2024)
Privacy Law Enforcement Under Centralized Governance: A Qualitative Analysis of Four Years' Special Privacy Rectification Campaigns
by: Jing, Tao, et al.
Published: (2025)
by: Jing, Tao, et al.
Published: (2025)
Play Guessing Game with LLM: Indirect Jailbreak Attack with Implicit Clues
by: Chang, Zhiyuan, et al.
Published: (2024)
by: Chang, Zhiyuan, et al.
Published: (2024)
AdaptAuth: Multi-Layered Behavioral and Credential Analysis for a Secure and Adaptive Authentication Framework for Password Security
by: Ghosh, Tonmoy
Published: (2025)
by: Ghosh, Tonmoy
Published: (2025)
Measure-Observe-Remeasure: An Interactive Paradigm for Differentially-Private Exploratory Analysis
by: Nanayakkara, Priyanka, et al.
Published: (2024)
by: Nanayakkara, Priyanka, et al.
Published: (2024)
Similar Items
-
Risk Psychology & Cyber-Attack Tactics
by: Kim, Rubens, et al.
Published: (2025) -
Professor X: Manipulating EEG BCI with Invisible and Robust Backdoor Attack
by: Liu, Xuan-Hao, et al.
Published: (2024) -
Benchmarking and Understanding Safety Risks in AI Character Platforms
by: Wei, Yiluo, et al.
Published: (2025) -
SoK: Come Together -- Unifying Security, Information Theory, and Cognition for a Mixed Reality Deception Attack Ontology & Analysis Framework
by: Teymourian, Ali, et al.
Published: (2025) -
Multiverse Privacy Theory for Contextual Risks in Complex User-AI Interactions
by: Gumusel, Ece
Published: (2025)