The Laminar Flow Hypothesis: Detecting Jailbreaks via Semantic Turbulence in Large Language Models
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Rahman, Md. Hasib Ur |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Improved Large Language Model Jailbreak Detection via Pretrained Embeddings
von: Galinkin, Erick, et al.
Veröffentlicht: (2024)
von: Galinkin, Erick, et al.
Veröffentlicht: (2024)
Reasoning-targeted Jailbreak Attacks on Large Reasoning Models via Semantic Triggers and Psychological Framing
von: Wang, Zehao, et al.
Veröffentlicht: (2026)
von: Wang, Zehao, et al.
Veröffentlicht: (2026)
Single-pass Detection of Jailbreaking Input in Large Language Models
von: Candogan, Leyla Naz, et al.
Veröffentlicht: (2025)
von: Candogan, Leyla Naz, et al.
Veröffentlicht: (2025)
Jailbreaking Black Box Large Language Models in Twenty Queries
von: Chao, Patrick, et al.
Veröffentlicht: (2023)
von: Chao, Patrick, et al.
Veröffentlicht: (2023)
Second-Order Information Matters: Revisiting Machine Unlearning for Large Language Models
von: Gu, Kang, et al.
Veröffentlicht: (2024)
von: Gu, Kang, et al.
Veröffentlicht: (2024)
BiasJailbreak:Analyzing Ethical Biases and Jailbreak Vulnerabilities in Large Language Models
von: Lee, Isack, et al.
Veröffentlicht: (2024)
von: Lee, Isack, et al.
Veröffentlicht: (2024)
SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks
von: Robey, Alexander, et al.
Veröffentlicht: (2023)
von: Robey, Alexander, et al.
Veröffentlicht: (2023)
Jailbreak Scaling Laws for Large Language Models: Polynomial-Exponential Crossover
von: Halder, Indranil, et al.
Veröffentlicht: (2026)
von: Halder, Indranil, et al.
Veröffentlicht: (2026)
Jailbreaking and Mitigation of Vulnerabilities in Large Language Models
von: Peng, Benji, et al.
Veröffentlicht: (2024)
von: Peng, Benji, et al.
Veröffentlicht: (2024)
The Consistency Hypothesis in Uncertainty Quantification for Large Language Models
von: Xiao, Quan, et al.
Veröffentlicht: (2025)
von: Xiao, Quan, et al.
Veröffentlicht: (2025)
Relative Positioning Based Code Chunking Method For Rich Context Retrieval In Repository Level Code Completion Task With Code Language Model
von: Rahman, Imranur, et al.
Veröffentlicht: (2025)
von: Rahman, Imranur, et al.
Veröffentlicht: (2025)
The Linear Representation Hypothesis and the Geometry of Large Language Models
von: Park, Kiho, et al.
Veröffentlicht: (2023)
von: Park, Kiho, et al.
Veröffentlicht: (2023)
How a Bit Becomes a Story: Semantic Steering via Differentiable Fault Injection
von: Haider, Zafaryab, et al.
Veröffentlicht: (2025)
von: Haider, Zafaryab, et al.
Veröffentlicht: (2025)
Grid2Guide: A* Enabled Small Language Model for Indoor Navigation
von: Haque, Md. Wasiul, et al.
Veröffentlicht: (2025)
von: Haque, Md. Wasiul, et al.
Veröffentlicht: (2025)
TROJail: Trajectory-Level Optimization for Multi-Turn Large Language Model Jailbreaks with Process Rewards
von: Xiong, Xiqiao, et al.
Veröffentlicht: (2025)
von: Xiong, Xiqiao, et al.
Veröffentlicht: (2025)
Hallucination to Truth: A Review of Fact-Checking and Factuality Evaluation in Large Language Models
von: Rahman, Subhey Sadi, et al.
Veröffentlicht: (2025)
von: Rahman, Subhey Sadi, et al.
Veröffentlicht: (2025)
Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring
von: Hua, Peichun, et al.
Veröffentlicht: (2025)
von: Hua, Peichun, et al.
Veröffentlicht: (2025)
Hypothesis Generation with Large Language Models
von: Zhou, Yangqiaoyu, et al.
Veröffentlicht: (2024)
von: Zhou, Yangqiaoyu, et al.
Veröffentlicht: (2024)
Jailbreaking Large Language Models with Symbolic Mathematics
von: Bethany, Emet, et al.
Veröffentlicht: (2024)
von: Bethany, Emet, et al.
Veröffentlicht: (2024)
Latent Semantic Manifolds in Large Language Models
von: Mabrok, Mohamed A.
Veröffentlicht: (2026)
von: Mabrok, Mohamed A.
Veröffentlicht: (2026)
AbFlowNet: Optimizing Antibody-Antigen Binding Energy via Diffusion-GFlowNet Fusion
von: Abir, Abrar Rahman, et al.
Veröffentlicht: (2025)
von: Abir, Abrar Rahman, et al.
Veröffentlicht: (2025)
EnJa: Ensemble Jailbreak on Large Language Models
von: Zhang, Jiahao, et al.
Veröffentlicht: (2024)
von: Zhang, Jiahao, et al.
Veröffentlicht: (2024)
LatentBreak: Jailbreaking Large Language Models through Latent Space Feedback
von: Mura, Raffaele, et al.
Veröffentlicht: (2025)
von: Mura, Raffaele, et al.
Veröffentlicht: (2025)
iCost: A Novel Instance Complexity Based Cost-Sensitive Learning Framework
von: Newaz, Asif, et al.
Veröffentlicht: (2024)
von: Newaz, Asif, et al.
Veröffentlicht: (2024)
SequentialBreak: Large Language Models Can be Fooled by Embedding Jailbreak Prompts into Sequential Prompt Chains
von: Saiem, Bijoy Ahmed, et al.
Veröffentlicht: (2024)
von: Saiem, Bijoy Ahmed, et al.
Veröffentlicht: (2024)
Gradient Cuff: Detecting Jailbreak Attacks on Large Language Models by Exploring Refusal Loss Landscapes
von: Hu, Xiaomeng, et al.
Veröffentlicht: (2024)
von: Hu, Xiaomeng, et al.
Veröffentlicht: (2024)
Permissive Information-Flow Analysis for Large Language Models
von: Siddiqui, Shoaib Ahmed, et al.
Veröffentlicht: (2024)
von: Siddiqui, Shoaib Ahmed, et al.
Veröffentlicht: (2024)
Embracing Large Language Models in Traffic Flow Forecasting
von: Zhao, Yusheng, et al.
Veröffentlicht: (2024)
von: Zhao, Yusheng, et al.
Veröffentlicht: (2024)
Open-RAG: Enhanced Retrieval-Augmented Reasoning with Open-Source Large Language Models
von: Islam, Shayekh Bin, et al.
Veröffentlicht: (2024)
von: Islam, Shayekh Bin, et al.
Veröffentlicht: (2024)
Emotion Detection From Social Media Posts
von: Rahman, Md Mahbubur, et al.
Veröffentlicht: (2023)
von: Rahman, Md Mahbubur, et al.
Veröffentlicht: (2023)
Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models
von: Ball, Sarah, et al.
Veröffentlicht: (2024)
von: Ball, Sarah, et al.
Veröffentlicht: (2024)
Towards Explainable Traffic Flow Prediction with Large Language Models
von: Guo, Xusen, et al.
Veröffentlicht: (2024)
von: Guo, Xusen, et al.
Veröffentlicht: (2024)
Improving Uncertainty Quantification in Large Language Models via Semantic Embeddings
von: Grewal, Yashvir S., et al.
Veröffentlicht: (2024)
von: Grewal, Yashvir S., et al.
Veröffentlicht: (2024)
Stop Taking Tokenizers for Granted: They Are Core Design Decisions in Large Language Models
von: Alqahtani, Sawsan, et al.
Veröffentlicht: (2026)
von: Alqahtani, Sawsan, et al.
Veröffentlicht: (2026)
An Explainable Transformer-based Model for Phishing Email Detection: A Large Language Model Approach
von: Uddin, Mohammad Amaz, et al.
Veröffentlicht: (2024)
von: Uddin, Mohammad Amaz, et al.
Veröffentlicht: (2024)
Zer0-Jack: A Memory-efficient Gradient-based Jailbreaking Method for Black-box Multi-modal Large Language Models
von: Chen, Tiejin, et al.
Veröffentlicht: (2024)
von: Chen, Tiejin, et al.
Veröffentlicht: (2024)
Jailbreak-AudioBench: In-Depth Evaluation and Analysis of Jailbreak Threats for Large Audio Language Models
von: Cheng, Hao, et al.
Veröffentlicht: (2025)
von: Cheng, Hao, et al.
Veröffentlicht: (2025)
Semantic-aware Wasserstein Policy Regularization for Large Language Model Alignment
von: Na, Byeonghu, et al.
Veröffentlicht: (2026)
von: Na, Byeonghu, et al.
Veröffentlicht: (2026)
EvoJail: Evolutionary Diverse Jailbreak Prompt Generation for Large Language Models
von: Tang, Rui, et al.
Veröffentlicht: (2026)
von: Tang, Rui, et al.
Veröffentlicht: (2026)
UniGuard: Towards Universal Safety Guardrails for Jailbreak Attacks on Multimodal Large Language Models
von: Oh, Sejoon, et al.
Veröffentlicht: (2024)
von: Oh, Sejoon, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Improved Large Language Model Jailbreak Detection via Pretrained Embeddings
von: Galinkin, Erick, et al.
Veröffentlicht: (2024) -
Reasoning-targeted Jailbreak Attacks on Large Reasoning Models via Semantic Triggers and Psychological Framing
von: Wang, Zehao, et al.
Veröffentlicht: (2026) -
Single-pass Detection of Jailbreaking Input in Large Language Models
von: Candogan, Leyla Naz, et al.
Veröffentlicht: (2025) -
Jailbreaking Black Box Large Language Models in Twenty Queries
von: Chao, Patrick, et al.
Veröffentlicht: (2023) -
Second-Order Information Matters: Revisiting Machine Unlearning for Large Language Models
von: Gu, Kang, et al.
Veröffentlicht: (2024)