Harnessing LLM to Attack LLM-Guarded Text-to-Image Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Deng, Yimo, Chen, Huangxun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Energy-Driven Adaptive Visual Token Pruning for Efficient Vision-Language Models
di: He, Jialuo, et al.
Pubblicazione: (2026)
di: He, Jialuo, et al.
Pubblicazione: (2026)
BlindGuard: Safeguarding LLM-based Multi-Agent Systems under Unknown Attacks
di: Miao, Rui, et al.
Pubblicazione: (2025)
di: Miao, Rui, et al.
Pubblicazione: (2025)
Towards Compact and Robust DNNs via Compression-aware Sharpness Minimization
di: He, Jialuo, et al.
Pubblicazione: (2026)
di: He, Jialuo, et al.
Pubblicazione: (2026)
RepreGuard: Detecting LLM-Generated Text by Revealing Hidden Representation Patterns
di: Chen, Xin, et al.
Pubblicazione: (2025)
di: Chen, Xin, et al.
Pubblicazione: (2025)
Harnessing LLM Agents with Skill Programs
di: Liu, Hongjun, et al.
Pubblicazione: (2026)
di: Liu, Hongjun, et al.
Pubblicazione: (2026)
StyleGuard: Preventing Text-to-Image-Model-based Style Mimicry Attacks by Style Perturbations
di: Li, Yanjie, et al.
Pubblicazione: (2025)
di: Li, Yanjie, et al.
Pubblicazione: (2025)
Reason2Attack: Jailbreaking Text-to-Image Models via LLM Reasoning
di: Zhang, Chenyu, et al.
Pubblicazione: (2025)
di: Zhang, Chenyu, et al.
Pubblicazione: (2025)
PersGuard: Preventing Malicious Personalization via Backdoor Attacks on Pre-trained Text-to-Image Diffusion Models
di: Liu, Xinwei, et al.
Pubblicazione: (2025)
di: Liu, Xinwei, et al.
Pubblicazione: (2025)
Autonomous Algorithm Discovery for Ptychography via Evolutionary LLM Reasoning
di: Yin, Xiangyu, et al.
Pubblicazione: (2026)
di: Yin, Xiangyu, et al.
Pubblicazione: (2026)
CodeGuard: Improving LLM Guardrails in CS Education
di: Raihan, Nishat, et al.
Pubblicazione: (2026)
di: Raihan, Nishat, et al.
Pubblicazione: (2026)
AutoBridge: Automating Smart Device Integration with Centralized Platform
di: Liu, Siyuan, et al.
Pubblicazione: (2025)
di: Liu, Siyuan, et al.
Pubblicazione: (2025)
BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks
di: Tu, Xinming, et al.
Pubblicazione: (2026)
di: Tu, Xinming, et al.
Pubblicazione: (2026)
Adapting the Interface, Not the Model: Runtime Harness Adaptation for Deterministic LLM Agents
di: Xu, Tianshi, et al.
Pubblicazione: (2026)
di: Xu, Tianshi, et al.
Pubblicazione: (2026)
GuardReasoner: Towards Reasoning-based LLM Safeguards
di: Liu, Yue, et al.
Pubblicazione: (2025)
di: Liu, Yue, et al.
Pubblicazione: (2025)
Harnessing LLM for Noise-Robust Cognitive Diagnosis in Web-Based Intelligent Education Systems
di: Zhang, Guixian, et al.
Pubblicazione: (2025)
di: Zhang, Guixian, et al.
Pubblicazione: (2025)
Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents
di: Lin, Minhua, et al.
Pubblicazione: (2026)
di: Lin, Minhua, et al.
Pubblicazione: (2026)
PrefixGuard: From LLM-Agent Traces to Online Failure-Warning Monitors
di: Huang, Xinmiao, et al.
Pubblicazione: (2026)
di: Huang, Xinmiao, et al.
Pubblicazione: (2026)
Distillation Traps and Guards: A Calibration Knob for LLM Distillability
di: Zhan, Weixiao, et al.
Pubblicazione: (2026)
di: Zhan, Weixiao, et al.
Pubblicazione: (2026)
ProbGuard: Probabilistic Runtime Monitoring for LLM Agent Safety
di: Wang, Haoyu, et al.
Pubblicazione: (2025)
di: Wang, Haoyu, et al.
Pubblicazione: (2025)
Privacy Guard & Token Parsimony by Prompt and Context Handling and LLM Routing
di: Langiu, Alessio
Pubblicazione: (2026)
di: Langiu, Alessio
Pubblicazione: (2026)
RouteGuard: Internal-Signal Detection of Skill Poisoning in LLM Agents
di: Xiao, Wenjie, et al.
Pubblicazione: (2026)
di: Xiao, Wenjie, et al.
Pubblicazione: (2026)
Judge Reliability Harness: Stress Testing the Reliability of LLM Judges
di: Dev, Sunishchal, et al.
Pubblicazione: (2026)
di: Dev, Sunishchal, et al.
Pubblicazione: (2026)
HARIVO: Harnessing Text-to-Image Models for Video Generation
di: Kwon, Mingi, et al.
Pubblicazione: (2024)
di: Kwon, Mingi, et al.
Pubblicazione: (2024)
AlignGuard: Scalable Safety Alignment for Text-to-Image Generation
di: Liu, Runtao, et al.
Pubblicazione: (2024)
di: Liu, Runtao, et al.
Pubblicazione: (2024)
CourtGuard: A Model-Agnostic Framework for Zero-Shot Policy Adaptation in LLM Safety
di: Suleymanov, Umid, et al.
Pubblicazione: (2026)
di: Suleymanov, Umid, et al.
Pubblicazione: (2026)
Semantic-level Backdoor Attack against Text-to-Image Diffusion Models
di: Chen, Tianxin, et al.
Pubblicazione: (2026)
di: Chen, Tianxin, et al.
Pubblicazione: (2026)
Your Language Model Can Secretly Write Like Humans: Contrastive Paraphrase Attacks on LLM-Generated Text Detectors
di: Fang, Hao, et al.
Pubblicazione: (2025)
di: Fang, Hao, et al.
Pubblicazione: (2025)
Gala: Global LLM Agents for Text-to-Model Translation
di: Cai, Junyang, et al.
Pubblicazione: (2025)
di: Cai, Junyang, et al.
Pubblicazione: (2025)
GroupGuard: A Framework for Modeling and Defending Collusive Attacks in Multi-Agent Systems
di: Tao, Yiling, et al.
Pubblicazione: (2026)
di: Tao, Yiling, et al.
Pubblicazione: (2026)
ASGuard: Activation-Scaling Guard to Mitigate Targeted Jailbreaking Attack
di: Park, Yein, et al.
Pubblicazione: (2025)
di: Park, Yein, et al.
Pubblicazione: (2025)
DistillGuard: Evaluating Defenses Against LLM Knowledge Distillation
di: Jiang, Bo
Pubblicazione: (2026)
di: Jiang, Bo
Pubblicazione: (2026)
Efficient Detection of LLM-generated Texts with a Bayesian Surrogate Model
di: Miao, Yibo, et al.
Pubblicazione: (2023)
di: Miao, Yibo, et al.
Pubblicazione: (2023)
MergeGuard: Efficient Thwarting of Trojan Attacks in Machine Learning Models
di: Shabgahi, Soheil Zibakhsh, et al.
Pubblicazione: (2025)
di: Shabgahi, Soheil Zibakhsh, et al.
Pubblicazione: (2025)
General Modular Harness for LLM Agents in Multi-Turn Gaming Environments
di: Zhang, Yuxuan, et al.
Pubblicazione: (2025)
di: Zhang, Yuxuan, et al.
Pubblicazione: (2025)
FlexGuard: Continuous Risk Scoring for Strictness-Adaptive LLM Content Moderation
di: Ding, Zhihao, et al.
Pubblicazione: (2026)
di: Ding, Zhihao, et al.
Pubblicazione: (2026)
QGuard:Question-based Zero-shot Guard for Multi-modal LLM Safety
di: Lee, Taegyeong, et al.
Pubblicazione: (2025)
di: Lee, Taegyeong, et al.
Pubblicazione: (2025)
Stop Comparing LLM Agents Without Disclosing the Harness
di: Zhang, Yunbei, et al.
Pubblicazione: (2026)
di: Zhang, Yunbei, et al.
Pubblicazione: (2026)
Harnessing Consistency for Robust Test-Time LLM Ensemble
di: Zeng, Zhichen, et al.
Pubblicazione: (2025)
di: Zeng, Zhichen, et al.
Pubblicazione: (2025)
CheatAgent: Attacking LLM-Empowered Recommender Systems via LLM Agent
di: Ning, Liang-bo, et al.
Pubblicazione: (2025)
di: Ning, Liang-bo, et al.
Pubblicazione: (2025)
LLM2: Let Large Language Models Harness System 2 Reasoning
di: Yang, Cheng, et al.
Pubblicazione: (2024)
di: Yang, Cheng, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Energy-Driven Adaptive Visual Token Pruning for Efficient Vision-Language Models
di: He, Jialuo, et al.
Pubblicazione: (2026) -
BlindGuard: Safeguarding LLM-based Multi-Agent Systems under Unknown Attacks
di: Miao, Rui, et al.
Pubblicazione: (2025) -
Towards Compact and Robust DNNs via Compression-aware Sharpness Minimization
di: He, Jialuo, et al.
Pubblicazione: (2026) -
RepreGuard: Detecting LLM-Generated Text by Revealing Hidden Representation Patterns
di: Chen, Xin, et al.
Pubblicazione: (2025) -
Harnessing LLM Agents with Skill Programs
di: Liu, Hongjun, et al.
Pubblicazione: (2026)