Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhuang, Jun, Jin, Haibo, Zhang, Ye, Kang, Zhengjian, Zhang, Wenbin, Dagher, Gaby G., Wang, Haohan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Jailbreaking Large Language Models Against Moderation Guardrails via Cipher Characters
von: Jin, Haibo, et al.
Veröffentlicht: (2024)
von: Jin, Haibo, et al.
Veröffentlicht: (2024)
Now You Hear Me: Audio Narrative Attacks Against Large Audio-Language Models
von: Yu, Ye, et al.
Veröffentlicht: (2026)
von: Yu, Ye, et al.
Veröffentlicht: (2026)
InfoFlood: Jailbreaking Large Language Models with Information Overload
von: Yadav, Advait, et al.
Veröffentlicht: (2025)
von: Yadav, Advait, et al.
Veröffentlicht: (2025)
Learning to Conceal Risk: Controllable Multi-turn Red Teaming for LLMs in the Financial Domain
von: Cheng, Gang, et al.
Veröffentlicht: (2025)
von: Cheng, Gang, et al.
Veröffentlicht: (2025)
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models
von: Jin, Haibo, et al.
Veröffentlicht: (2025)
von: Jin, Haibo, et al.
Veröffentlicht: (2025)
GuardVal: Dynamic Large Language Model Jailbreak Evaluation for Comprehensive Safety Testing
von: Zhang, Peiyan, et al.
Veröffentlicht: (2025)
von: Zhang, Peiyan, et al.
Veröffentlicht: (2025)
JailbreakZoo: Survey, Landscapes, and Horizons in Jailbreaking Large Language and Vision-Language Models
von: Jin, Haibo, et al.
Veröffentlicht: (2024)
von: Jin, Haibo, et al.
Veröffentlicht: (2024)
Reasoning Can Hurt the Inductive Abilities of Large Language Models
von: Jin, Haibo, et al.
Veröffentlicht: (2025)
von: Jin, Haibo, et al.
Veröffentlicht: (2025)
LoRA-Guard: Parameter-Efficient Guardrail Adaptation for Content Moderation of Large Language Models
von: Elesedy, Hayder, et al.
Veröffentlicht: (2024)
von: Elesedy, Hayder, et al.
Veröffentlicht: (2024)
Probing Causality Manipulation of Large Language Models
von: Zhang, Chenyang, et al.
Veröffentlicht: (2024)
von: Zhang, Chenyang, et al.
Veröffentlicht: (2024)
Blockchain for Large Language Model Security and Safety: A Holistic Survey
von: Geren, Caleb, et al.
Veröffentlicht: (2024)
von: Geren, Caleb, et al.
Veröffentlicht: (2024)
GUARD: Role-playing to Generate Natural-language Jailbreakings to Test Guideline Adherence of Large Language Models
von: Jin, Haibo, et al.
Veröffentlicht: (2024)
von: Jin, Haibo, et al.
Veröffentlicht: (2024)
Building Guardrails for Large Language Models
von: Dong, Yi, et al.
Veröffentlicht: (2024)
von: Dong, Yi, et al.
Veröffentlicht: (2024)
Virus: Harmful Fine-tuning Attack for Large Language Models Bypassing Guardrail Moderation
von: Huang, Tiansheng, et al.
Veröffentlicht: (2025)
von: Huang, Tiansheng, et al.
Veröffentlicht: (2025)
The Unappreciated Role of Intent in Algorithmic Moderation of Social Media Content
von: Wang, Xinyu, et al.
Veröffentlicht: (2024)
von: Wang, Xinyu, et al.
Veröffentlicht: (2024)
RigorLLM: Resilient Guardrails for Large Language Models against Undesired Content
von: Yuan, Zhuowen, et al.
Veröffentlicht: (2024)
von: Yuan, Zhuowen, et al.
Veröffentlicht: (2024)
OpenGuardrails: A Configurable, Unified, and Scalable Guardrails Platform for Large Language Models
von: Wang, Thomas, et al.
Veröffentlicht: (2025)
von: Wang, Thomas, et al.
Veröffentlicht: (2025)
A Causal Explainable Guardrails for Large Language Models
von: Chu, Zhixuan, et al.
Veröffentlicht: (2024)
von: Chu, Zhixuan, et al.
Veröffentlicht: (2024)
Adaptive Honeypot Allocation in Multi-Attacker Networks via Bayesian Stackelberg Games
von: Park, Dongyoung, et al.
Veröffentlicht: (2025)
von: Park, Dongyoung, et al.
Veröffentlicht: (2025)
SGuard-v1: Safety Guardrail for Large Language Models
von: Lee, JoonHo, et al.
Veröffentlicht: (2025)
von: Lee, JoonHo, et al.
Veröffentlicht: (2025)
Legilimens: Practical and Unified Content Moderation for Large Language Model Services
von: Wu, Jialin, et al.
Veröffentlicht: (2024)
von: Wu, Jialin, et al.
Veröffentlicht: (2024)
A Non-Zero-Sum Game Model for Optimal Cyber Defense Strategies
von: Park, Dongyoung, et al.
Veröffentlicht: (2025)
von: Park, Dongyoung, et al.
Veröffentlicht: (2025)
SafeRoute: Adaptive Model Selection for Efficient and Accurate Safety Guardrails in Large Language Models
von: Lee, Seanie, et al.
Veröffentlicht: (2025)
von: Lee, Seanie, et al.
Veröffentlicht: (2025)
Trust-Oriented Adaptive Guardrails for Large Language Models
von: Hu, Jinwei, et al.
Veröffentlicht: (2024)
von: Hu, Jinwei, et al.
Veröffentlicht: (2024)
Wi-Chat: Large Language Model Powered Wi-Fi Sensing
von: Zhang, Haopeng, et al.
Veröffentlicht: (2025)
von: Zhang, Haopeng, et al.
Veröffentlicht: (2025)
Learning to Communicate: Toward End-to-End Optimization of Multi-Agent Language Systems
von: Yu, Ye, et al.
Veröffentlicht: (2026)
von: Yu, Ye, et al.
Veröffentlicht: (2026)
Break Me If You Can: Self-Jailbreaking of Aligned LLMs via Lexical Insertion Prompting
von: Kulshreshtha, Devang, et al.
Veröffentlicht: (2026)
von: Kulshreshtha, Devang, et al.
Veröffentlicht: (2026)
SentGuard: Sentence-Level Streaming Guardrails for Large Language Models
von: Yu, Jiaqi, et al.
Veröffentlicht: (2026)
von: Yu, Jiaqi, et al.
Veröffentlicht: (2026)
ChallengeMe: An Adversarial Learning-enabled Text Summarization Framework
von: Deng, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Deng, Xiaoyu, et al.
Veröffentlicht: (2025)
IntentGPT: Few-shot Intent Discovery with Large Language Models
von: Rodriguez, Juan A., et al.
Veröffentlicht: (2024)
von: Rodriguez, Juan A., et al.
Veröffentlicht: (2024)
Agent Primitives: Reusable Latent Building Blocks for Multi-Agent Systems
von: Jin, Haibo, et al.
Veröffentlicht: (2026)
von: Jin, Haibo, et al.
Veröffentlicht: (2026)
Semi-Supervised Learning for Large Language Models Safety and Content Moderation
von: Dinuta, Eduard Stefan, et al.
Veröffentlicht: (2025)
von: Dinuta, Eduard Stefan, et al.
Veröffentlicht: (2025)
Fake News Detection and Manipulation Reasoning via Large Vision-Language Models
von: Jin, Ruihan, et al.
Veröffentlicht: (2024)
von: Jin, Ruihan, et al.
Veröffentlicht: (2024)
Unlocking the Future: Exploring Look-Ahead Planning Mechanistic Interpretability in Large Language Models
von: Men, Tianyi, et al.
Veröffentlicht: (2024)
von: Men, Tianyi, et al.
Veröffentlicht: (2024)
Query Performance Explanation through Large Language Model for HTAP Systems
von: Xiu, Haibo, et al.
Veröffentlicht: (2024)
von: Xiu, Haibo, et al.
Veröffentlicht: (2024)
Large Language Models are Vulnerable to Bait-and-Switch Attacks for Generating Harmful Content
von: Bianchi, Federico, et al.
Veröffentlicht: (2024)
von: Bianchi, Federico, et al.
Veröffentlicht: (2024)
TextReg: Mitigating Prompt Distributional Overfitting via Regularized Text-Space Optimization
von: Fu, Lucheng, et al.
Veröffentlicht: (2026)
von: Fu, Lucheng, et al.
Veröffentlicht: (2026)
Fairness in Large Language Models: A Taxonomic Survey
von: Chu, Zhibo, et al.
Veröffentlicht: (2024)
von: Chu, Zhibo, et al.
Veröffentlicht: (2024)
Watch Your Language: Investigating Content Moderation with Large Language Models
von: Kumar, Deepak, et al.
Veröffentlicht: (2023)
von: Kumar, Deepak, et al.
Veröffentlicht: (2023)
Capsule Network-Based Semantic Intent Modeling for Human-Computer Interaction
von: Wang, Shixiao, et al.
Veröffentlicht: (2025)
von: Wang, Shixiao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Jailbreaking Large Language Models Against Moderation Guardrails via Cipher Characters
von: Jin, Haibo, et al.
Veröffentlicht: (2024) -
Now You Hear Me: Audio Narrative Attacks Against Large Audio-Language Models
von: Yu, Ye, et al.
Veröffentlicht: (2026) -
InfoFlood: Jailbreaking Large Language Models with Information Overload
von: Yadav, Advait, et al.
Veröffentlicht: (2025) -
Learning to Conceal Risk: Controllable Multi-turn Red Teaming for LLMs in the Financial Domain
von: Cheng, Gang, et al.
Veröffentlicht: (2025) -
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models
von: Jin, Haibo, et al.
Veröffentlicht: (2025)