Harnessing LLM to Attack LLM-Guarded Text-to-Image Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Deng, Yimo, Chen, Huangxun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Energy-Driven Adaptive Visual Token Pruning for Efficient Vision-Language Models
von: He, Jialuo, et al.
Veröffentlicht: (2026)
von: He, Jialuo, et al.
Veröffentlicht: (2026)
BlindGuard: Safeguarding LLM-based Multi-Agent Systems under Unknown Attacks
von: Miao, Rui, et al.
Veröffentlicht: (2025)
von: Miao, Rui, et al.
Veröffentlicht: (2025)
Towards Compact and Robust DNNs via Compression-aware Sharpness Minimization
von: He, Jialuo, et al.
Veröffentlicht: (2026)
von: He, Jialuo, et al.
Veröffentlicht: (2026)
RepreGuard: Detecting LLM-Generated Text by Revealing Hidden Representation Patterns
von: Chen, Xin, et al.
Veröffentlicht: (2025)
von: Chen, Xin, et al.
Veröffentlicht: (2025)
Harnessing LLM Agents with Skill Programs
von: Liu, Hongjun, et al.
Veröffentlicht: (2026)
von: Liu, Hongjun, et al.
Veröffentlicht: (2026)
StyleGuard: Preventing Text-to-Image-Model-based Style Mimicry Attacks by Style Perturbations
von: Li, Yanjie, et al.
Veröffentlicht: (2025)
von: Li, Yanjie, et al.
Veröffentlicht: (2025)
Reason2Attack: Jailbreaking Text-to-Image Models via LLM Reasoning
von: Zhang, Chenyu, et al.
Veröffentlicht: (2025)
von: Zhang, Chenyu, et al.
Veröffentlicht: (2025)
PersGuard: Preventing Malicious Personalization via Backdoor Attacks on Pre-trained Text-to-Image Diffusion Models
von: Liu, Xinwei, et al.
Veröffentlicht: (2025)
von: Liu, Xinwei, et al.
Veröffentlicht: (2025)
Autonomous Algorithm Discovery for Ptychography via Evolutionary LLM Reasoning
von: Yin, Xiangyu, et al.
Veröffentlicht: (2026)
von: Yin, Xiangyu, et al.
Veröffentlicht: (2026)
CodeGuard: Improving LLM Guardrails in CS Education
von: Raihan, Nishat, et al.
Veröffentlicht: (2026)
von: Raihan, Nishat, et al.
Veröffentlicht: (2026)
AutoBridge: Automating Smart Device Integration with Centralized Platform
von: Liu, Siyuan, et al.
Veröffentlicht: (2025)
von: Liu, Siyuan, et al.
Veröffentlicht: (2025)
BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks
von: Tu, Xinming, et al.
Veröffentlicht: (2026)
von: Tu, Xinming, et al.
Veröffentlicht: (2026)
Adapting the Interface, Not the Model: Runtime Harness Adaptation for Deterministic LLM Agents
von: Xu, Tianshi, et al.
Veröffentlicht: (2026)
von: Xu, Tianshi, et al.
Veröffentlicht: (2026)
GuardReasoner: Towards Reasoning-based LLM Safeguards
von: Liu, Yue, et al.
Veröffentlicht: (2025)
von: Liu, Yue, et al.
Veröffentlicht: (2025)
Harnessing LLM for Noise-Robust Cognitive Diagnosis in Web-Based Intelligent Education Systems
von: Zhang, Guixian, et al.
Veröffentlicht: (2025)
von: Zhang, Guixian, et al.
Veröffentlicht: (2025)
Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents
von: Lin, Minhua, et al.
Veröffentlicht: (2026)
von: Lin, Minhua, et al.
Veröffentlicht: (2026)
PrefixGuard: From LLM-Agent Traces to Online Failure-Warning Monitors
von: Huang, Xinmiao, et al.
Veröffentlicht: (2026)
von: Huang, Xinmiao, et al.
Veröffentlicht: (2026)
Distillation Traps and Guards: A Calibration Knob for LLM Distillability
von: Zhan, Weixiao, et al.
Veröffentlicht: (2026)
von: Zhan, Weixiao, et al.
Veröffentlicht: (2026)
ProbGuard: Probabilistic Runtime Monitoring for LLM Agent Safety
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
Privacy Guard & Token Parsimony by Prompt and Context Handling and LLM Routing
von: Langiu, Alessio
Veröffentlicht: (2026)
von: Langiu, Alessio
Veröffentlicht: (2026)
RouteGuard: Internal-Signal Detection of Skill Poisoning in LLM Agents
von: Xiao, Wenjie, et al.
Veröffentlicht: (2026)
von: Xiao, Wenjie, et al.
Veröffentlicht: (2026)
Judge Reliability Harness: Stress Testing the Reliability of LLM Judges
von: Dev, Sunishchal, et al.
Veröffentlicht: (2026)
von: Dev, Sunishchal, et al.
Veröffentlicht: (2026)
HARIVO: Harnessing Text-to-Image Models for Video Generation
von: Kwon, Mingi, et al.
Veröffentlicht: (2024)
von: Kwon, Mingi, et al.
Veröffentlicht: (2024)
AlignGuard: Scalable Safety Alignment for Text-to-Image Generation
von: Liu, Runtao, et al.
Veröffentlicht: (2024)
von: Liu, Runtao, et al.
Veröffentlicht: (2024)
CourtGuard: A Model-Agnostic Framework for Zero-Shot Policy Adaptation in LLM Safety
von: Suleymanov, Umid, et al.
Veröffentlicht: (2026)
von: Suleymanov, Umid, et al.
Veröffentlicht: (2026)
Semantic-level Backdoor Attack against Text-to-Image Diffusion Models
von: Chen, Tianxin, et al.
Veröffentlicht: (2026)
von: Chen, Tianxin, et al.
Veröffentlicht: (2026)
Your Language Model Can Secretly Write Like Humans: Contrastive Paraphrase Attacks on LLM-Generated Text Detectors
von: Fang, Hao, et al.
Veröffentlicht: (2025)
von: Fang, Hao, et al.
Veröffentlicht: (2025)
Gala: Global LLM Agents for Text-to-Model Translation
von: Cai, Junyang, et al.
Veröffentlicht: (2025)
von: Cai, Junyang, et al.
Veröffentlicht: (2025)
GroupGuard: A Framework for Modeling and Defending Collusive Attacks in Multi-Agent Systems
von: Tao, Yiling, et al.
Veröffentlicht: (2026)
von: Tao, Yiling, et al.
Veröffentlicht: (2026)
ASGuard: Activation-Scaling Guard to Mitigate Targeted Jailbreaking Attack
von: Park, Yein, et al.
Veröffentlicht: (2025)
von: Park, Yein, et al.
Veröffentlicht: (2025)
DistillGuard: Evaluating Defenses Against LLM Knowledge Distillation
von: Jiang, Bo
Veröffentlicht: (2026)
von: Jiang, Bo
Veröffentlicht: (2026)
Efficient Detection of LLM-generated Texts with a Bayesian Surrogate Model
von: Miao, Yibo, et al.
Veröffentlicht: (2023)
von: Miao, Yibo, et al.
Veröffentlicht: (2023)
MergeGuard: Efficient Thwarting of Trojan Attacks in Machine Learning Models
von: Shabgahi, Soheil Zibakhsh, et al.
Veröffentlicht: (2025)
von: Shabgahi, Soheil Zibakhsh, et al.
Veröffentlicht: (2025)
General Modular Harness for LLM Agents in Multi-Turn Gaming Environments
von: Zhang, Yuxuan, et al.
Veröffentlicht: (2025)
von: Zhang, Yuxuan, et al.
Veröffentlicht: (2025)
FlexGuard: Continuous Risk Scoring for Strictness-Adaptive LLM Content Moderation
von: Ding, Zhihao, et al.
Veröffentlicht: (2026)
von: Ding, Zhihao, et al.
Veröffentlicht: (2026)
QGuard:Question-based Zero-shot Guard for Multi-modal LLM Safety
von: Lee, Taegyeong, et al.
Veröffentlicht: (2025)
von: Lee, Taegyeong, et al.
Veröffentlicht: (2025)
Stop Comparing LLM Agents Without Disclosing the Harness
von: Zhang, Yunbei, et al.
Veröffentlicht: (2026)
von: Zhang, Yunbei, et al.
Veröffentlicht: (2026)
Harnessing Consistency for Robust Test-Time LLM Ensemble
von: Zeng, Zhichen, et al.
Veröffentlicht: (2025)
von: Zeng, Zhichen, et al.
Veröffentlicht: (2025)
CheatAgent: Attacking LLM-Empowered Recommender Systems via LLM Agent
von: Ning, Liang-bo, et al.
Veröffentlicht: (2025)
von: Ning, Liang-bo, et al.
Veröffentlicht: (2025)
LLM2: Let Large Language Models Harness System 2 Reasoning
von: Yang, Cheng, et al.
Veröffentlicht: (2024)
von: Yang, Cheng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Energy-Driven Adaptive Visual Token Pruning for Efficient Vision-Language Models
von: He, Jialuo, et al.
Veröffentlicht: (2026) -
BlindGuard: Safeguarding LLM-based Multi-Agent Systems under Unknown Attacks
von: Miao, Rui, et al.
Veröffentlicht: (2025) -
Towards Compact and Robust DNNs via Compression-aware Sharpness Minimization
von: He, Jialuo, et al.
Veröffentlicht: (2026) -
RepreGuard: Detecting LLM-Generated Text by Revealing Hidden Representation Patterns
von: Chen, Xin, et al.
Veröffentlicht: (2025) -
Harnessing LLM Agents with Skill Programs
von: Liu, Hongjun, et al.
Veröffentlicht: (2026)