JailbreakLens: Visual Analysis of Jailbreak Attacks Against Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Feng, Yingchaojie, Chen, Zhizhang, Kang, Zhining, Wang, Sijia, Tian, Haoyu, Zhang, Wei, Zhu, Minfeng, Chen, Wei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Play Guessing Game with LLM: Indirect Jailbreak Attack with Implicit Clues
by: Chang, Zhiyuan, et al.
Published: (2024)
by: Chang, Zhiyuan, et al.
Published: (2024)
JailbreakLens: Interpreting Jailbreak Mechanism in the Lens of Representation and Circuit
by: He, Zeqing, et al.
Published: (2024)
by: He, Zeqing, et al.
Published: (2024)
Hacc-Man: An Arcade Game for Jailbreaking LLMs
by: Valentim, Matheus, et al.
Published: (2024)
by: Valentim, Matheus, et al.
Published: (2024)
RAGExplorer: A Visual Analytics System for the Comparative Diagnosis of RAG Systems
by: Tian, Haoyu, et al.
Published: (2026)
by: Tian, Haoyu, et al.
Published: (2026)
Professor X: Manipulating EEG BCI with Invisible and Robust Backdoor Attack
by: Liu, Xuan-Hao, et al.
Published: (2024)
by: Liu, Xuan-Hao, et al.
Published: (2024)
Language Model Agents Under Attack: A Cross Model-Benchmark of Profit-Seeking Behaviors in Customer Service
by: Zhang, Jingyu
Published: (2025)
by: Zhang, Jingyu
Published: (2025)
Risk Psychology & Cyber-Attack Tactics
by: Kim, Rubens, et al.
Published: (2025)
by: Kim, Rubens, et al.
Published: (2025)
RLHFPoison: Reward Poisoning Attack for Reinforcement Learning with Human Feedback in Large Language Models
by: Wang, Jiongxiao, et al.
Published: (2023)
by: Wang, Jiongxiao, et al.
Published: (2023)
Listen to the Voices of Everyday Users: Democratizing Privacy Ratings for Sensitive Data Access in Mobile Apps
by: Wang, Liu, et al.
Published: (2026)
by: Wang, Liu, et al.
Published: (2026)
"Tab, Tab, Bug": Security Pitfalls of Next Edit Suggestions in AI-Integrated IDEs
by: Lyu, Yunlong, et al.
Published: (2026)
by: Lyu, Yunlong, et al.
Published: (2026)
Hidden-in-Plain-Text: A Benchmark for Social-Web Indirect Prompt Injection in RAG
by: Guo, Haoze, et al.
Published: (2026)
by: Guo, Haoze, et al.
Published: (2026)
SoK: The Privacy Paradox of Large Language Models: Advancements, Privacy Risks, and Mitigation
by: Shanmugarasa, Yashothara, et al.
Published: (2025)
by: Shanmugarasa, Yashothara, et al.
Published: (2025)
InjectLab: A Tactical Framework for Adversarial Threat Modeling Against Large Language Models
by: Howard, Austin
Published: (2025)
by: Howard, Austin
Published: (2025)
Understanding User Privacy Perceptions of GenAI Smartphones
by: Jin, Ran, et al.
Published: (2026)
by: Jin, Ran, et al.
Published: (2026)
Comparative Simulation of Phishing Attacks on a Critical Information Infrastructure Organization: An Empirical Study
by: Sirawongphatsara, Patsita, et al.
Published: (2024)
by: Sirawongphatsara, Patsita, et al.
Published: (2024)
Can Large Language Models Automate Phishing Warning Explanations? A Controlled Experiment on Effectiveness and User Perception
by: Cau, Federico Maria, et al.
Published: (2025)
by: Cau, Federico Maria, et al.
Published: (2025)
From Perception to Protection: A Developer-Centered Study of Security and Privacy Threats in Extended Reality (XR)
by: Cai, Kunlin, et al.
Published: (2025)
by: Cai, Kunlin, et al.
Published: (2025)
"Did They F***ing Consent to That?": Safer Digital Intimacy via Proactive Protection Against Image-Based Sexual Abuse
by: Qin, Lucy, et al.
Published: (2024)
by: Qin, Lucy, et al.
Published: (2024)
Synopticon: Consensus-Based Cheating Detection System for Competitive Games
by: Kang, Jeuk, et al.
Published: (2025)
by: Kang, Jeuk, et al.
Published: (2025)
SoK: Come Together -- Unifying Security, Information Theory, and Cognition for a Mixed Reality Deception Attack Ontology & Analysis Framework
by: Teymourian, Ali, et al.
Published: (2025)
by: Teymourian, Ali, et al.
Published: (2025)
BioMoTouch: Touch-Based Behavioral Authentication via Biometric-Motion Interaction Modeling
by: Ling, Zijian, et al.
Published: (2026)
by: Ling, Zijian, et al.
Published: (2026)
Invisible, Unreadable, and Inaudible Cookie Notices: An Evaluation of Cookie Notices for Users with Visual Impairments
by: Clarke, James M., et al.
Published: (2023)
by: Clarke, James M., et al.
Published: (2023)
Defogger: A Visual Analysis Approach for Data Exploration of Sensitive Data Protected by Differential Privacy
by: Wang, Xumeng, et al.
Published: (2024)
by: Wang, Xumeng, et al.
Published: (2024)
Buck You: Designing Easy-to-Onboard Blockchain Applications with Zero-Knowledge Login and Sponsored Transactions on Sui
by: Chen, Eason, et al.
Published: (2024)
by: Chen, Eason, et al.
Published: (2024)
Anti-Sensing: Defense against Unauthorized Radar-based Human Vital Sign Sensing with Physically Realizable Wearable Oscillators
by: Oshim, Md Farhan Tasnim, et al.
Published: (2025)
by: Oshim, Md Farhan Tasnim, et al.
Published: (2025)
Unveiling Privacy and Security Gaps in Female Health Apps
by: Hassan, Muhammad, et al.
Published: (2025)
by: Hassan, Muhammad, et al.
Published: (2025)
Towards Proactive Defense Against Cyber Cognitive Attacks
by: Rushing, Bonnie, et al.
Published: (2025)
by: Rushing, Bonnie, et al.
Published: (2025)
Toward Accessible Mobile Money: A Voice-Driven, Biometrically Secured USSD Automation Framework for Visually Impaired Users
by: Ajayi, Sunday, et al.
Published: (2026)
by: Ajayi, Sunday, et al.
Published: (2026)
A Survey of Wireless Sensing Security from a Role-Based View: Victim, Weapon, and Shield
by: Geng, Ruixu, et al.
Published: (2024)
by: Geng, Ruixu, et al.
Published: (2024)
Towards Scalable Defenses against Intimate Partner Infiltrations
by: Yang, Weisi, et al.
Published: (2025)
by: Yang, Weisi, et al.
Published: (2025)
Evaluating the Usability of Differential Privacy Tools with Data Practitioners
by: Ngong, Ivoline C., et al.
Published: (2023)
by: Ngong, Ivoline C., et al.
Published: (2023)
CultiVerse: Towards Cross-Cultural Understanding for Paintings with Large Language Model
by: Zhang, Wei, et al.
Published: (2024)
by: Zhang, Wei, et al.
Published: (2024)
Benchmarking and Understanding Safety Risks in AI Character Platforms
by: Wei, Yiluo, et al.
Published: (2025)
by: Wei, Yiluo, et al.
Published: (2025)
False Reality: Uncovering Sensor-induced Human-VR Interaction Vulnerability
by: Jiang, Yancheng, et al.
Published: (2025)
by: Jiang, Yancheng, et al.
Published: (2025)
Anti-Phishing Training (Still) Does Not Work: A Large-Scale Reproduction of Phishing Training Inefficacy Grounded in the NIST Phish Scale
by: Rozema, Andrew T., et al.
Published: (2025)
by: Rozema, Andrew T., et al.
Published: (2025)
SoK: Usability Studies in Differential Privacy
by: Dibia, Onyinye, et al.
Published: (2024)
by: Dibia, Onyinye, et al.
Published: (2024)
SECURE: Benchmarking Large Language Models for Cybersecurity
by: Bhusal, Dipkamal, et al.
Published: (2024)
by: Bhusal, Dipkamal, et al.
Published: (2024)
SoK (or SoLK?): On the Quantitative Study of Sociodemographic Factors and Computer Security Behaviors
by: Wei, Miranda, et al.
Published: (2024)
by: Wei, Miranda, et al.
Published: (2024)
Human-Centered Privacy Research in the Age of Large Language Models
by: Li, Tianshi, et al.
Published: (2024)
by: Li, Tianshi, et al.
Published: (2024)
Adversarial Attacks on Machine Learning-Aided Visualizations
by: Fujiwara, Takanori, et al.
Published: (2024)
by: Fujiwara, Takanori, et al.
Published: (2024)
Similar Items
-
Play Guessing Game with LLM: Indirect Jailbreak Attack with Implicit Clues
by: Chang, Zhiyuan, et al.
Published: (2024) -
JailbreakLens: Interpreting Jailbreak Mechanism in the Lens of Representation and Circuit
by: He, Zeqing, et al.
Published: (2024) -
Hacc-Man: An Arcade Game for Jailbreaking LLMs
by: Valentim, Matheus, et al.
Published: (2024) -
RAGExplorer: A Visual Analytics System for the Comparative Diagnosis of RAG Systems
by: Tian, Haoyu, et al.
Published: (2026) -
Professor X: Manipulating EEG BCI with Invisible and Robust Backdoor Attack
by: Liu, Xuan-Hao, et al.
Published: (2024)