SpatialJB: How Text Distribution Art Becomes the "Jailbreak Key" for LLM Guardrails
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mou, Zhiyi, Yang, Jingyuan, Qian, Zeheng, Ni, Wangze, Xiao, Tianfang, Liu, Ning, Zhang, Chen, Qin, Zhan, Ren, Kui |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SRBench: A Comprehensive Benchmark for Sequential Recommendation with Large Language Models
von: Li, Jianhong, et al.
Veröffentlicht: (2026)
von: Li, Jianhong, et al.
Veröffentlicht: (2026)
DualBreach: Efficient Dual-Jailbreaking via Target-Driven Initialization and Multi-Target Optimization
von: Huang, Xinzhe, et al.
Veröffentlicht: (2025)
von: Huang, Xinzhe, et al.
Veröffentlicht: (2025)
QCFuse: Query-Centric Cache Fusion for Efficient RAG Inference
von: Yan, Jianxin, et al.
Veröffentlicht: (2026)
von: Yan, Jianxin, et al.
Veröffentlicht: (2026)
ContextCache: Context-Aware Semantic Cache for Multi-Turn Queries in Large Language Models
von: Yan, Jianxin, et al.
Veröffentlicht: (2025)
von: Yan, Jianxin, et al.
Veröffentlicht: (2025)
ERASER: Machine Unlearning in MLaaS via an Inference Serving-Aware Approach
von: Hu, Yuke, et al.
Veröffentlicht: (2023)
von: Hu, Yuke, et al.
Veröffentlicht: (2023)
HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
von: Yang, Langqi, et al.
Veröffentlicht: (2025)
von: Yang, Langqi, et al.
Veröffentlicht: (2025)
A Causal Explainable Guardrails for Large Language Models
von: Chu, Zhixuan, et al.
Veröffentlicht: (2024)
von: Chu, Zhixuan, et al.
Veröffentlicht: (2024)
JANUS: A Lightweight Framework for Jailbreaking Text-to-Image Models via Distribution Optimization
von: Zheng, Haolun, et al.
Veröffentlicht: (2026)
von: Zheng, Haolun, et al.
Veröffentlicht: (2026)
Untargeted Jailbreak Attack
von: Huang, Xinzhe, et al.
Veröffentlicht: (2025)
von: Huang, Xinzhe, et al.
Veröffentlicht: (2025)
Numerical Estimation of Spatial Distributions under Differential Privacy
von: Du, Leilei, et al.
Veröffentlicht: (2024)
von: Du, Leilei, et al.
Veröffentlicht: (2024)
Active Honeypot Guardrail System: Probing and Confirming Multi-Turn LLM Jailbreaks
von: Wu, ChenYu, et al.
Veröffentlicht: (2025)
von: Wu, ChenYu, et al.
Veröffentlicht: (2025)
Eguard: Defending LLM Embeddings Against Inversion Attacks via Text Mutual Information Optimization
von: Liu, Tiantian, et al.
Veröffentlicht: (2024)
von: Liu, Tiantian, et al.
Veröffentlicht: (2024)
MAJIC: Markovian Adaptive Jailbreaking via Iterative Composition of Diverse Innovative Strategies
von: Qi, Weiwei, et al.
Veröffentlicht: (2025)
von: Qi, Weiwei, et al.
Veröffentlicht: (2025)
TRACE: Task-Aware Adaptive Self-Evolving Agentic Jailbreaking
von: Zeng, Churui, et al.
Veröffentlicht: (2026)
von: Zeng, Churui, et al.
Veröffentlicht: (2026)
The Art of the Jailbreak: Formulating Jailbreak Attacks for LLM Security Beyond Binary Scoring
von: Hossain, Ismail, et al.
Veröffentlicht: (2026)
von: Hossain, Ismail, et al.
Veröffentlicht: (2026)
Channel-Level Semantic Perturbations: Unlearnable Examples for Diverse Training Paradigms
von: Wang, Bo, et al.
Veröffentlicht: (2026)
von: Wang, Bo, et al.
Veröffentlicht: (2026)
Don't Say No: Jailbreaking LLM by Suppressing Refusal
von: Zhou, Yukai, et al.
Veröffentlicht: (2024)
von: Zhou, Yukai, et al.
Veröffentlicht: (2024)
Guarding the Guardrails: A Taxonomy-Driven Approach to Jailbreak Detection
von: Giarrusso, Francesco, et al.
Veröffentlicht: (2025)
von: Giarrusso, Francesco, et al.
Veröffentlicht: (2025)
SoK: Evaluating Jailbreak Guardrails for Large Language Models
von: Wang, Xunguang, et al.
Veröffentlicht: (2025)
von: Wang, Xunguang, et al.
Veröffentlicht: (2025)
The Art of Becoming Infinite
von: Stanchina, Gabriella
Veröffentlicht: (2025)
von: Stanchina, Gabriella
Veröffentlicht: (2025)
Forewarned is Forearmed: Pre-Synthesizing Jailbreak-like Instructions to Enhance LLM Safety Guardrail to Potential Attacks
von: Liu, Sheng, et al.
Veröffentlicht: (2025)
von: Liu, Sheng, et al.
Veröffentlicht: (2025)
Bypassing LLM Guardrails: An Empirical Analysis of Evasion Attacks against Prompt Injection and Jailbreak Detection Systems
von: Hackett, William, et al.
Veröffentlicht: (2025)
von: Hackett, William, et al.
Veröffentlicht: (2025)
ConsisGuard: Aligning Safety Deliberation with Policy Enforcement in LLM Guardrails
von: Wang, Yan, et al.
Veröffentlicht: (2026)
von: Wang, Yan, et al.
Veröffentlicht: (2026)
Distributionally Robust Policy Learning under Concept Drifts
von: Wang, Jingyuan, et al.
Veröffentlicht: (2024)
von: Wang, Jingyuan, et al.
Veröffentlicht: (2024)
OneShield -- the Next Generation of LLM Guardrails
von: DeLuca, Chad, et al.
Veröffentlicht: (2025)
von: DeLuca, Chad, et al.
Veröffentlicht: (2025)
RAC: Relation-Aware Cache Replacement for Large Language Models
von: Wu, Yuchong, et al.
Veröffentlicht: (2026)
von: Wu, Yuchong, et al.
Veröffentlicht: (2026)
StructRide: A Framework to Exploit the Structure Information of Shareability Graph in Ridesharing
von: Zhan, Jiexi, et al.
Veröffentlicht: (2024)
von: Zhan, Jiexi, et al.
Veröffentlicht: (2024)
Jailbreaking Large Language Models Against Moderation Guardrails via Cipher Characters
von: Jin, Haibo, et al.
Veröffentlicht: (2024)
von: Jin, Haibo, et al.
Veröffentlicht: (2024)
Silencing the Guardrails: Inference-Time Jailbreaking via Dynamic Contextual Representation Ablation
von: Xing, Wenpeng, et al.
Veröffentlicht: (2026)
von: Xing, Wenpeng, et al.
Veröffentlicht: (2026)
Sequential Gaussian Avatars with Hierarchical Motion Context
von: Xu, Wangze, et al.
Veröffentlicht: (2024)
von: Xu, Wangze, et al.
Veröffentlicht: (2024)
ConceptGuard: Neuro-Symbolic Safety Guardrails via Sparse Interpretable Jailbreak Concepts
von: Aswal, Darpan, et al.
Veröffentlicht: (2025)
von: Aswal, Darpan, et al.
Veröffentlicht: (2025)
Reclaiming Idle CPU Cycles on Kubernetes: Sparse-Domain Multiplexing for Concurrent MPI-CFD Simulations
von: Xie, Tianfang
Veröffentlicht: (2026)
von: Xie, Tianfang
Veröffentlicht: (2026)
Rank-Aware Resource Scheduling for Tightly-Coupled MPI Workloads on Kubernetes
von: Xie, Tianfang
Veröffentlicht: (2026)
von: Xie, Tianfang
Veröffentlicht: (2026)
Certified Minimax Unlearning with Generalization Rates and Deletion Capacity
von: Liu, Jiaqi, et al.
Veröffentlicht: (2023)
von: Liu, Jiaqi, et al.
Veröffentlicht: (2023)
Prompt-Consistency Image Generation (PCIG): A Unified Framework Integrating LLMs, Knowledge Graphs, and Controllable Diffusion Models
von: Sun, Yichen, et al.
Veröffentlicht: (2024)
von: Sun, Yichen, et al.
Veröffentlicht: (2024)
FINER: Enhancing State-of-the-art Classifiers with Feature Attribution to Facilitate Security Analysis
von: He, Yiling, et al.
Veröffentlicht: (2023)
von: He, Yiling, et al.
Veröffentlicht: (2023)
Mapping the Microenvironment: How Spatial‐Omics Is Advancing the Understanding of Allograft Fate
von: Lisha Mou, et al.
Veröffentlicht: (2025)
von: Lisha Mou, et al.
Veröffentlicht: (2025)
Towards Evaluation for Real-World LLM Unlearning
von: Miao, Ke, et al.
Veröffentlicht: (2025)
von: Miao, Ke, et al.
Veröffentlicht: (2025)
Safety Guardrails for LLM-Enabled Robots
von: Ravichandran, Zachary, et al.
Veröffentlicht: (2025)
von: Ravichandran, Zachary, et al.
Veröffentlicht: (2025)
Sora Detector: A Unified Hallucination Detection for Large Text-to-Video Models
von: Chu, Zhixuan, et al.
Veröffentlicht: (2024)
von: Chu, Zhixuan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
SRBench: A Comprehensive Benchmark for Sequential Recommendation with Large Language Models
von: Li, Jianhong, et al.
Veröffentlicht: (2026) -
DualBreach: Efficient Dual-Jailbreaking via Target-Driven Initialization and Multi-Target Optimization
von: Huang, Xinzhe, et al.
Veröffentlicht: (2025) -
QCFuse: Query-Centric Cache Fusion for Efficient RAG Inference
von: Yan, Jianxin, et al.
Veröffentlicht: (2026) -
ContextCache: Context-Aware Semantic Cache for Multi-Turn Queries in Large Language Models
von: Yan, Jianxin, et al.
Veröffentlicht: (2025) -
ERASER: Machine Unlearning in MLaaS via an Inference Serving-Aware Approach
von: Hu, Yuke, et al.
Veröffentlicht: (2023)