Red-Teaming Text-to-Image Models via In-Context Experience Replay and Semantic-Preserving Prompt Rewriting
Fuente:
arXiv
Guardado en:
| Autores principales: | Chin, Zhi-Yi, Chen, Pin-Yu, Chiu, Wei-Chen, Fritz, Mario |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Be Your Own Red Teamer: Safety Alignment via Self-Play and Reflective Experience Replay
por: Wang, Hao, et al.
Publicado: (2026)
por: Wang, Hao, et al.
Publicado: (2026)
GenBreak: Red Teaming Text-to-Image Generators Using Large Language Models
por: Wang, Zilong, et al.
Publicado: (2025)
por: Wang, Zilong, et al.
Publicado: (2025)
Prompt Optimization and Evaluation for LLM Automated Red Teaming
por: Freenor, Michael, et al.
Publicado: (2025)
por: Freenor, Michael, et al.
Publicado: (2025)
Automated Progressive Red Teaming
por: Jiang, Bojian, et al.
Publicado: (2024)
por: Jiang, Bojian, et al.
Publicado: (2024)
ContextualJailbreak: Evolutionary Red-Teaming via Simulated Conversational Priming
por: Béjar, Mario Rodríguez, et al.
Publicado: (2026)
por: Béjar, Mario Rodríguez, et al.
Publicado: (2026)
Prompting4Debugging: Red-Teaming Text-to-Image Diffusion Models by Finding Problematic Prompts
por: Chin, Zhi-Yi, et al.
Publicado: (2023)
por: Chin, Zhi-Yi, et al.
Publicado: (2023)
RedTWIZ: Diverse LLM Red Teaming via Adaptive Attack Planning
por: Horal, Artur, et al.
Publicado: (2025)
por: Horal, Artur, et al.
Publicado: (2025)
Red Teaming the Mind of the Machine: A Systematic Evaluation of Prompt Injection and Jailbreak Vulnerabilities in LLMs
por: Pathade, Chetan
Publicado: (2025)
por: Pathade, Chetan
Publicado: (2025)
Red Teaming AI Red Teaming
por: Majumdar, Subhabrata, et al.
Publicado: (2025)
por: Majumdar, Subhabrata, et al.
Publicado: (2025)
DREAM: Scalable Red Teaming for Text-to-Image Generative Systems via Distribution Modeling
por: Li, Boheng, et al.
Publicado: (2025)
por: Li, Boheng, et al.
Publicado: (2025)
Learning-Based Automated Adversarial Red-Teaming for Robustness Evaluation of Large Language Models
por: Wei, Zhang, et al.
Publicado: (2025)
por: Wei, Zhang, et al.
Publicado: (2025)
RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent
por: Xu, Huiyu, et al.
Publicado: (2024)
por: Xu, Huiyu, et al.
Publicado: (2024)
DP-BART for Privatized Text Rewriting under Local Differential Privacy
por: Igamberdiev, Timour, et al.
Publicado: (2023)
por: Igamberdiev, Timour, et al.
Publicado: (2023)
AEIOU: A Unified Defense Framework against NSFW Prompts in Text-to-Image Models
por: Wang, Yiming, et al.
Publicado: (2024)
por: Wang, Yiming, et al.
Publicado: (2024)
Resource Consumption Red-Teaming for Large Vision-Language Models
por: Gao, Haoran, et al.
Publicado: (2025)
por: Gao, Haoran, et al.
Publicado: (2025)
Training a General Purpose Automated Red Teaming Model
por: Padmakumar, Aishwarya, et al.
Publicado: (2026)
por: Padmakumar, Aishwarya, et al.
Publicado: (2026)
Safe Text-to-Image Generation: Simply Sanitize the Prompt Embedding
por: Qiu, Huming, et al.
Publicado: (2024)
por: Qiu, Huming, et al.
Publicado: (2024)
Value-Aligned Prompt Moderation via Zero-Shot Agentic Rewriting for Safe Image Generation
por: Zhao, Xin, et al.
Publicado: (2025)
por: Zhao, Xin, et al.
Publicado: (2025)
MIRAGE: Multimodal Immersive Reasoning and Guided Exploration for Red-Team Jailbreak Attacks
por: You, Wenhao, et al.
Publicado: (2025)
por: You, Wenhao, et al.
Publicado: (2025)
Spend Your Budget Wisely: Towards an Intelligent Distribution of the Privacy Budget in Differentially Private Text Rewriting
por: Meisenbacher, Stephen, et al.
Publicado: (2025)
por: Meisenbacher, Stephen, et al.
Publicado: (2025)
Beyond Theoretical Bounds: Empirical Privacy Loss Calibration for Text Rewriting Under Local Differential Privacy
por: Li, Weijun, et al.
Publicado: (2026)
por: Li, Weijun, et al.
Publicado: (2026)
Red-Teaming Text-to-Image Systems by Rule-based Preference Modeling
por: Cao, Yichuan, et al.
Publicado: (2025)
por: Cao, Yichuan, et al.
Publicado: (2025)
SafeSearch: Automated Red-Teaming of LLM-Based Search Agents
por: Dong, Jianshuo, et al.
Publicado: (2025)
por: Dong, Jianshuo, et al.
Publicado: (2025)
Rethinking and Red-Teaming Protective Perturbation in Personalized Diffusion Models
por: Liu, Yixin, et al.
Publicado: (2024)
por: Liu, Yixin, et al.
Publicado: (2024)
OpenRT: An Open-Source Red Teaming Framework for Multimodal LLMs
por: Wang, Xin, et al.
Publicado: (2026)
por: Wang, Xin, et al.
Publicado: (2026)
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs
por: Hu, Xiaomeng, et al.
Publicado: (2025)
por: Hu, Xiaomeng, et al.
Publicado: (2025)
SurrogatePrompt: Bypassing the Safety Filter of Text-to-Image Models via Substitution
por: Ba, Zhongjie, et al.
Publicado: (2023)
por: Ba, Zhongjie, et al.
Publicado: (2023)
Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models
por: Hu, Xiaomeng, et al.
Publicado: (2024)
por: Hu, Xiaomeng, et al.
Publicado: (2024)
SIRAJ: Diverse and Efficient Red-Teaming for LLM Agents via Distilled Structured Reasoning
por: Zhou, Kaiwen, et al.
Publicado: (2025)
por: Zhou, Kaiwen, et al.
Publicado: (2025)
From Coordinates to Context: An LLM-Bootstrapped Semantic Encoding Framework for Privacy-Preserving Mobile Sensing Stress Recognition
por: Phan, Hoang Khang, et al.
Publicado: (2025)
por: Phan, Hoang Khang, et al.
Publicado: (2025)
Adaptive Instruction Composition for Automated LLM Red-Teaming
por: Zymet, Jesse, et al.
Publicado: (2026)
por: Zymet, Jesse, et al.
Publicado: (2026)
Jailbreak-Zero: A Path to Pareto Optimal Red Teaming for Large Language Models
por: Hu, Kai, et al.
Publicado: (2025)
por: Hu, Kai, et al.
Publicado: (2025)
Adversarial Nibbler: An Open Red-Teaming Method for Identifying Diverse Harms in Text-to-Image Generation
por: Quaye, Jessica, et al.
Publicado: (2024)
por: Quaye, Jessica, et al.
Publicado: (2024)
Red Teaming GPT-4V: Are GPT-4V Safe Against Uni/Multi-Modal Jailbreak Attacks?
por: Chen, Shuo, et al.
Publicado: (2024)
por: Chen, Shuo, et al.
Publicado: (2024)
Beyond Text Prompts: Precise Concept Erasure through Text-Image Collaboration
por: Li, Jun, et al.
Publicado: (2026)
por: Li, Jun, et al.
Publicado: (2026)
Summon a Demon and Bind it: A Grounded Theory of LLM Red Teaming
por: Inie, Nanna, et al.
Publicado: (2023)
por: Inie, Nanna, et al.
Publicado: (2023)
Groot: Adversarial Testing for Generative Text-to-Image Models with Tree-based Semantic Transformation
por: Liu, Yi, et al.
Publicado: (2024)
por: Liu, Yi, et al.
Publicado: (2024)
Defensive Prompt Patch: A Robust and Interpretable Defense of LLMs against Jailbreak Attacks
por: Xiong, Chen, et al.
Publicado: (2024)
por: Xiong, Chen, et al.
Publicado: (2024)
FlashRT: Towards Computationally and Memory Efficient Red-Teaming for Prompt Injection and Knowledge Corruption
por: Wang, Yanting, et al.
Publicado: (2026)
por: Wang, Yanting, et al.
Publicado: (2026)
RedTeamLLM: an Agentic AI framework for offensive security
por: Challita, Brian, et al.
Publicado: (2025)
por: Challita, Brian, et al.
Publicado: (2025)
Ejemplares similares
-
Be Your Own Red Teamer: Safety Alignment via Self-Play and Reflective Experience Replay
por: Wang, Hao, et al.
Publicado: (2026) -
GenBreak: Red Teaming Text-to-Image Generators Using Large Language Models
por: Wang, Zilong, et al.
Publicado: (2025) -
Prompt Optimization and Evaluation for LLM Automated Red Teaming
por: Freenor, Michael, et al.
Publicado: (2025) -
Automated Progressive Red Teaming
por: Jiang, Bojian, et al.
Publicado: (2024) -
ContextualJailbreak: Evolutionary Red-Teaming via Simulated Conversational Priming
por: Béjar, Mario Rodríguez, et al.
Publicado: (2026)