Play Guessing Game with LLM: Indirect Jailbreak Attack with Implicit Clues
Fuente:
arXiv
Salvato in:
| Autori principali: | Chang, Zhiyuan, Li, Mingyang, Liu, Yi, Wang, Junjie, Wang, Qing, Liu, Yang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Hacc-Man: An Arcade Game for Jailbreaking LLMs
di: Valentim, Matheus, et al.
Pubblicazione: (2024)
di: Valentim, Matheus, et al.
Pubblicazione: (2024)
"It's a Fair Game", or Is It? Examining How Users Navigate Disclosure Risks and Benefits When Using LLM-Based Conversational Agents
di: Zhang, Zhiping, et al.
Pubblicazione: (2023)
di: Zhang, Zhiping, et al.
Pubblicazione: (2023)
From Assistants to Adversaries: Exploring the Security Risks of Mobile LLM Agents
di: Wu, Liangxuan, et al.
Pubblicazione: (2025)
di: Wu, Liangxuan, et al.
Pubblicazione: (2025)
LLM Whisperer: An Inconspicuous Attack to Bias LLM Responses
di: Lin, Weiran, et al.
Pubblicazione: (2024)
di: Lin, Weiran, et al.
Pubblicazione: (2024)
Autonomy Reshapes How Personalization Affects Privacy Concerns and Trust in LLM Agents
di: Zhang, Zhiping, et al.
Pubblicazione: (2025)
di: Zhang, Zhiping, et al.
Pubblicazione: (2025)
Mimicking the Familiar: Dynamic Command Generation for Information Theft Attacks in LLM Tool-Learning System
di: Jiang, Ziyou, et al.
Pubblicazione: (2025)
di: Jiang, Ziyou, et al.
Pubblicazione: (2025)
Rescriber: Smaller-LLM-Powered User-Led Data Minimization for LLM-Based Chatbots
di: Zhou, Jijie, et al.
Pubblicazione: (2024)
di: Zhou, Jijie, et al.
Pubblicazione: (2024)
Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework
di: Dassanayake, Rishane, et al.
Pubblicazione: (2025)
di: Dassanayake, Rishane, et al.
Pubblicazione: (2025)
PrivateXR: Defending Privacy Attacks in Extended Reality Through Explainable AI-Guided Differential Privacy
di: Kundu, Ripan Kumar, et al.
Pubblicazione: (2025)
di: Kundu, Ripan Kumar, et al.
Pubblicazione: (2025)
Cyri: A Conversational AI-based Assistant for Supporting the Human User in Detecting and Responding to Phishing Attacks
di: La Torre, Antonio, et al.
Pubblicazione: (2025)
di: La Torre, Antonio, et al.
Pubblicazione: (2025)
Current state of LLM Risks and AI Guardrails
di: Ayyamperumal, Suriya Ganesh, et al.
Pubblicazione: (2024)
di: Ayyamperumal, Suriya Ganesh, et al.
Pubblicazione: (2024)
One Shot Dominance: Knowledge Poisoning Attack on Retrieval-Augmented Generation Systems
di: Chang, Zhiyuan, et al.
Pubblicazione: (2025)
di: Chang, Zhiyuan, et al.
Pubblicazione: (2025)
Empowering Users in Digital Privacy Management through Interactive LLM-Based Agents
di: Sun, Bolun, et al.
Pubblicazione: (2024)
di: Sun, Bolun, et al.
Pubblicazione: (2024)
Adversarial Attacks on Machine Learning-Aided Visualizations
di: Fujiwara, Takanori, et al.
Pubblicazione: (2024)
di: Fujiwara, Takanori, et al.
Pubblicazione: (2024)
Human-Centered Privacy Research in the Age of Large Language Models
di: Li, Tianshi, et al.
Pubblicazione: (2024)
di: Li, Tianshi, et al.
Pubblicazione: (2024)
BounTCHA: A CAPTCHA Utilizing Boundary Identification in Guided Generative AI-extended Videos
di: Lin, Lehao, et al.
Pubblicazione: (2025)
di: Lin, Lehao, et al.
Pubblicazione: (2025)
JailbreakLens: Visual Analysis of Jailbreak Attacks Against Large Language Models
di: Feng, Yingchaojie, et al.
Pubblicazione: (2024)
di: Feng, Yingchaojie, et al.
Pubblicazione: (2024)
RLHFPoison: Reward Poisoning Attack for Reinforcement Learning with Human Feedback in Large Language Models
di: Wang, Jiongxiao, et al.
Pubblicazione: (2023)
di: Wang, Jiongxiao, et al.
Pubblicazione: (2023)
Privacy Leakage Overshadowed by Views of AI: A Study on Human Oversight of Privacy in Language Model Agent
di: Zhang, Zhiping, et al.
Pubblicazione: (2024)
di: Zhang, Zhiping, et al.
Pubblicazione: (2024)
This PIN Can Be Easily Guessed: Analyzing the Security of Smartphone Unlock PINs
di: Markert, Philipp, et al.
Pubblicazione: (2020)
di: Markert, Philipp, et al.
Pubblicazione: (2020)
"Impressively Scary:" Exploring User Perceptions and Reactions to Unraveling Machine Learning Models in Social Media Applications
di: West, Jack, et al.
Pubblicazione: (2025)
di: West, Jack, et al.
Pubblicazione: (2025)
Securing Virtual Reality Experiences: Unveiling and Tackling Cybersickness Attacks with Explainable AI
di: Kundu, Ripan Kumar, et al.
Pubblicazione: (2025)
di: Kundu, Ripan Kumar, et al.
Pubblicazione: (2025)
Learned, Lagged, LLM-splained: LLM Responses to End User Security Questions
di: Prakash, Vijay, et al.
Pubblicazione: (2024)
di: Prakash, Vijay, et al.
Pubblicazione: (2024)
Adversarial VR: An Open-Source Testbed for Evaluating Adversarial Robustness of VR Cybersickness Detection and Mitigation
di: Ahmed, Istiak, et al.
Pubblicazione: (2025)
di: Ahmed, Istiak, et al.
Pubblicazione: (2025)
Agentic AI and the Industrialization of Cyber Offense: Forecast, Consequences, and Defensive Priorities for Enterprises and the Mittelstand
di: Koch, Christopher
Pubblicazione: (2026)
di: Koch, Christopher
Pubblicazione: (2026)
Decision-Aware Trust Signal Alignment for SOC Alert Triage
di: Chowdhury, Israt Jahan, et al.
Pubblicazione: (2026)
di: Chowdhury, Israt Jahan, et al.
Pubblicazione: (2026)
MeAJOR Corpus: A Multi-Source Dataset for Phishing Email Detection
di: Mendes, Paulo, et al.
Pubblicazione: (2025)
di: Mendes, Paulo, et al.
Pubblicazione: (2025)
Towards Secure AI-driven Industrial Metaverse with NFT Digital Twins
di: Prakash, Ravi, et al.
Pubblicazione: (2024)
di: Prakash, Ravi, et al.
Pubblicazione: (2024)
JEEVHITAA -- An End-to-End HCAI System to Support Collective Care
di: Srinivasan, Shyama Sastha Krishnamoorthy, et al.
Pubblicazione: (2025)
di: Srinivasan, Shyama Sastha Krishnamoorthy, et al.
Pubblicazione: (2025)
Human-AI Collaboration in Cloud Security: Cognitive Hierarchy-Driven Deep Reinforcement Learning
di: Aref, Zahra, et al.
Pubblicazione: (2025)
di: Aref, Zahra, et al.
Pubblicazione: (2025)
AI-Assisted Adaptive Rendering for High-Frequency Security Telemetry in Web Interfaces
di: Rajhans, Mona
Pubblicazione: (2026)
di: Rajhans, Mona
Pubblicazione: (2026)
SECURE: Benchmarking Large Language Models for Cybersecurity
di: Bhusal, Dipkamal, et al.
Pubblicazione: (2024)
di: Bhusal, Dipkamal, et al.
Pubblicazione: (2024)
InjectLab: A Tactical Framework for Adversarial Threat Modeling Against Large Language Models
di: Howard, Austin
Pubblicazione: (2025)
di: Howard, Austin
Pubblicazione: (2025)
Human-Centered Explainability in AI-Enhanced UI Security Interfaces: Designing Trustworthy Copilots for Cybersecurity Analysts
di: Rajhans, Mona
Pubblicazione: (2026)
di: Rajhans, Mona
Pubblicazione: (2026)
Personalised Feedback Framework for Online Education Programmes Using Generative AI
di: Kuzminykh, Ievgeniia, et al.
Pubblicazione: (2024)
di: Kuzminykh, Ievgeniia, et al.
Pubblicazione: (2024)
Identify As A Human Does: A Pathfinder of Next-Generation Anti-Cheat Framework for First-Person Shooter Games
di: Zhang, Jiayi, et al.
Pubblicazione: (2024)
di: Zhang, Jiayi, et al.
Pubblicazione: (2024)
"Are You Sure?": An Empirical Study of Human Perception Vulnerability in LLM-Driven Agentic Systems
di: Li, Xinfeng, et al.
Pubblicazione: (2026)
di: Li, Xinfeng, et al.
Pubblicazione: (2026)
Known By Their Actions: Fingerprinting LLM Browser Agents via UI Traces
di: Lugoloobi, William, et al.
Pubblicazione: (2026)
di: Lugoloobi, William, et al.
Pubblicazione: (2026)
Malicious LLM-Based Conversational AI Makes Users Reveal Personal Information
di: Zhan, Xiao, et al.
Pubblicazione: (2025)
di: Zhan, Xiao, et al.
Pubblicazione: (2025)
AdInject: Real-World Black-Box Attacks on Web Agents via Advertising Delivery
di: Wang, Haowei, et al.
Pubblicazione: (2025)
di: Wang, Haowei, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Hacc-Man: An Arcade Game for Jailbreaking LLMs
di: Valentim, Matheus, et al.
Pubblicazione: (2024) -
"It's a Fair Game", or Is It? Examining How Users Navigate Disclosure Risks and Benefits When Using LLM-Based Conversational Agents
di: Zhang, Zhiping, et al.
Pubblicazione: (2023) -
From Assistants to Adversaries: Exploring the Security Risks of Mobile LLM Agents
di: Wu, Liangxuan, et al.
Pubblicazione: (2025) -
LLM Whisperer: An Inconspicuous Attack to Bias LLM Responses
di: Lin, Weiran, et al.
Pubblicazione: (2024) -
Autonomy Reshapes How Personalization Affects Privacy Concerns and Trust in LLM Agents
di: Zhang, Zhiping, et al.
Pubblicazione: (2025)