JADES: A Universal Framework for Jailbreak Assessment via Decompositional Scoring
Fuente:
arXiv
Salvato in:
| Autori principali: | Chu, Junjie, Li, Mingjie, Yang, Ziqing, Leng, Ye, Lin, Chenhao, Shen, Chao, Backes, Michael, Shen, Yun, Zhang, Yang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
di: Chu, Junjie, et al.
Pubblicazione: (2024)
di: Chu, Junjie, et al.
Pubblicazione: (2024)
When Understanding Becomes a Risk: Authenticity and Safety Risks in the Emerging Image Generation Paradigm
di: Leng, Ye, et al.
Pubblicazione: (2026)
di: Leng, Ye, et al.
Pubblicazione: (2026)
Benchmark of Benchmarks: Unpacking Influence and Code Repository Quality in LLM Safety Benchmarks
di: Chu, Junjie, et al.
Pubblicazione: (2026)
di: Chu, Junjie, et al.
Pubblicazione: (2026)
Understanding LLM Behavior When Encountering User-Supplied Harmful Content in Harmless Tasks
di: Chu, Junjie, et al.
Pubblicazione: (2026)
di: Chu, Junjie, et al.
Pubblicazione: (2026)
Adjacent Words, Divergent Intents: Jailbreaking Large Language Models via Task Concurrency
di: Jiang, Yukun, et al.
Pubblicazione: (2025)
di: Jiang, Yukun, et al.
Pubblicazione: (2025)
"Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
di: Shen, Xinyue, et al.
Pubblicazione: (2023)
di: Shen, Xinyue, et al.
Pubblicazione: (2023)
When GPT Spills the Tea: Comprehensive Assessment of Knowledge File Leakage in GPTs
di: Shen, Xinyue, et al.
Pubblicazione: (2025)
di: Shen, Xinyue, et al.
Pubblicazione: (2025)
Voice Jailbreak Attacks Against GPT-4o
di: Shen, Xinyue, et al.
Pubblicazione: (2024)
di: Shen, Xinyue, et al.
Pubblicazione: (2024)
Synthetic Artifact Auditing: Tracing LLM-Generated Synthetic Data Usage in Downstream Applications
di: Wu, Yixin, et al.
Pubblicazione: (2025)
di: Wu, Yixin, et al.
Pubblicazione: (2025)
The Challenge of Identifying the Origin of Black-Box Large Language Models
di: Yang, Ziqing, et al.
Pubblicazione: (2025)
di: Yang, Ziqing, et al.
Pubblicazione: (2025)
Reconstruct Your Previous Conversations! Comprehensively Investigating Privacy Leakage Risks in Conversations with GPT Models
di: Chu, Junjie, et al.
Pubblicazione: (2024)
di: Chu, Junjie, et al.
Pubblicazione: (2024)
HarmfulSkillBench: How Do Harmful Skills Weaponize Your Agents?
di: Jiang, Yukun, et al.
Pubblicazione: (2026)
di: Jiang, Yukun, et al.
Pubblicazione: (2026)
GEO-Detective: Unveiling Location Privacy Risks in Images with LLM Agents
di: Zhang, Xinyu, et al.
Pubblicazione: (2025)
di: Zhang, Xinyu, et al.
Pubblicazione: (2025)
JailPO: A Novel Black-box Jailbreak Framework via Preference Optimization against Aligned LLMs
di: Li, Hongyi, et al.
Pubblicazione: (2024)
di: Li, Hongyi, et al.
Pubblicazione: (2024)
"To Survive, I Must Defect": Jailbreaking LLMs via the Game-Theory Scenarios
di: Sun, Zhen, et al.
Pubblicazione: (2025)
di: Sun, Zhen, et al.
Pubblicazione: (2025)
BadBone: Backdoor Attacks Against Backbone Models in Visual Prompt Learning
di: Yang, Ziqing, et al.
Pubblicazione: (2026)
di: Yang, Ziqing, et al.
Pubblicazione: (2026)
Revisiting Training-Inference Trigger Intensity in Backdoor Attacks
di: Lin, Chenhao, et al.
Pubblicazione: (2025)
di: Lin, Chenhao, et al.
Pubblicazione: (2025)
Peering Behind the Shield: Guardrail Identification in Large Language Models
di: Yang, Ziqing, et al.
Pubblicazione: (2025)
di: Yang, Ziqing, et al.
Pubblicazione: (2025)
On the Proactive Generation of Unsafe Images From Text-To-Image Models Using Benign Prompts
di: Wu, Yixin, et al.
Pubblicazione: (2023)
di: Wu, Yixin, et al.
Pubblicazione: (2023)
Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities
di: Qu, Yiting, et al.
Pubblicazione: (2025)
di: Qu, Yiting, et al.
Pubblicazione: (2025)
Image-Perfect Imperfections: Safety, Bias, and Authenticity in the Shadow of Text-To-Image Model Evolution
di: Wu, Yixin, et al.
Pubblicazione: (2024)
di: Wu, Yixin, et al.
Pubblicazione: (2024)
Sequential Comics for Jailbreaking Multimodal Large Language Models via Structured Visual Storytelling
di: Zhang, Deyue, et al.
Pubblicazione: (2025)
di: Zhang, Deyue, et al.
Pubblicazione: (2025)
Sparse Models, Sparse Safety: Unsafe Routes in Mixture-of-Experts LLMs
di: Jiang, Yukun, et al.
Pubblicazione: (2026)
di: Jiang, Yukun, et al.
Pubblicazione: (2026)
Generated Data with Fake Privacy: Hidden Dangers of Fine-tuning Large Language Models on Generated Data
di: Akkus, Atilla, et al.
Pubblicazione: (2024)
di: Akkus, Atilla, et al.
Pubblicazione: (2024)
Enhancing Jailbreak Attacks on LLMs via Persona Prompts
di: Zhang, Zheng, et al.
Pubblicazione: (2025)
di: Zhang, Zheng, et al.
Pubblicazione: (2025)
TEMPLATEFUZZ: Fine-Grained Chat Template Fuzzing for Jailbreaking and Red Teaming LLMs
di: Shen, Qingchao, et al.
Pubblicazione: (2026)
di: Shen, Qingchao, et al.
Pubblicazione: (2026)
Transferable & Stealthy Ensemble Attacks: A Black-Box Jailbreaking Framework for Large Language Models
di: Yang, Yiqi, et al.
Pubblicazione: (2024)
di: Yang, Yiqi, et al.
Pubblicazione: (2024)
CodeChameleon: Personalized Encryption Framework for Jailbreaking Large Language Models
di: Lv, Huijie, et al.
Pubblicazione: (2024)
di: Lv, Huijie, et al.
Pubblicazione: (2024)
Robustness Over Time: Understanding Adversarial Examples' Effectiveness on Longitudinal Versions of Large Language Models
di: Liu, Yugeng, et al.
Pubblicazione: (2023)
di: Liu, Yugeng, et al.
Pubblicazione: (2023)
"Humans welcome to observe": A First Look at the Agent Social Network Moltbook
di: Jiang, Yukun, et al.
Pubblicazione: (2026)
di: Jiang, Yukun, et al.
Pubblicazione: (2026)
Bidirectional Intention Inference Enhances LLMs' Defense Against Multi-Turn Jailbreak Attacks
di: Tong, Haibo, et al.
Pubblicazione: (2025)
di: Tong, Haibo, et al.
Pubblicazione: (2025)
Real Money, Fake Models: Deceptive Model Claims in Shadow APIs
di: Zhang, Yage, et al.
Pubblicazione: (2026)
di: Zhang, Yage, et al.
Pubblicazione: (2026)
Guaranteed Jailbreaking Defense via Disrupt-and-Rectify Smoothing
di: Lin, Zheng, et al.
Pubblicazione: (2026)
di: Lin, Zheng, et al.
Pubblicazione: (2026)
The Art of the Jailbreak: Formulating Jailbreak Attacks for LLM Security Beyond Binary Scoring
di: Hossain, Ismail, et al.
Pubblicazione: (2026)
di: Hossain, Ismail, et al.
Pubblicazione: (2026)
DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs
di: Xu, Wenzhuo, et al.
Pubblicazione: (2026)
di: Xu, Wenzhuo, et al.
Pubblicazione: (2026)
Jailbreaking Attacks vs. Content Safety Filters: How Far Are We in the LLM Safety Arms Race?
di: Xin, Yuan, et al.
Pubblicazione: (2025)
di: Xin, Yuan, et al.
Pubblicazione: (2025)
SOS! Soft Prompt Attack Against Open-Source Large Language Models
di: Yang, Ziqing, et al.
Pubblicazione: (2024)
di: Yang, Ziqing, et al.
Pubblicazione: (2024)
Distract Large Language Models for Automatic Jailbreak Attack
di: Xiao, Zeguan, et al.
Pubblicazione: (2024)
di: Xiao, Zeguan, et al.
Pubblicazione: (2024)
Obscure but Effective: Classical Chinese Jailbreak Prompt Optimization via Bio-Inspired Search
di: Huang, Xun, et al.
Pubblicazione: (2026)
di: Huang, Xun, et al.
Pubblicazione: (2026)
Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models
di: Ma, Jiachen, et al.
Pubblicazione: (2024)
di: Ma, Jiachen, et al.
Pubblicazione: (2024)
Documenti analoghi
-
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
di: Chu, Junjie, et al.
Pubblicazione: (2024) -
When Understanding Becomes a Risk: Authenticity and Safety Risks in the Emerging Image Generation Paradigm
di: Leng, Ye, et al.
Pubblicazione: (2026) -
Benchmark of Benchmarks: Unpacking Influence and Code Repository Quality in LLM Safety Benchmarks
di: Chu, Junjie, et al.
Pubblicazione: (2026) -
Understanding LLM Behavior When Encountering User-Supplied Harmful Content in Harmless Tasks
di: Chu, Junjie, et al.
Pubblicazione: (2026) -
Adjacent Words, Divergent Intents: Jailbreaking Large Language Models via Task Concurrency
di: Jiang, Yukun, et al.
Pubblicazione: (2025)