RedTWIZ: Diverse LLM Red Teaming via Adaptive Attack Planning
Fuente:
arXiv
Saved in:
| Main Authors: | Horal, Artur, Pina, Daniel, Paz, Henrique, Paulo, Iago, Soares, João, Ferreira, Rafael, Tavares, Diogo, Glória-Silva, Diogo, Magalhães, João, Semedo, David |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TWIZ-v2: The Wizard of Multimodal Conversational-Stimulus
by: Ferreira, Rafael, et al.
Published: (2023)
by: Ferreira, Rafael, et al.
Published: (2023)
Plan-Grounded Large Language Models for Dual Goal Conversational Settings
by: Glória-Silva, Diogo, et al.
Published: (2024)
by: Glória-Silva, Diogo, et al.
Published: (2024)
Show and Guide: Instructional-Plan Grounded Vision and Language Model
by: Glória-Silva, Diogo, et al.
Published: (2024)
by: Glória-Silva, Diogo, et al.
Published: (2024)
ALBA: A European Portuguese Benchmark for Evaluating Language and Linguistic Dimensions in Generative LLMs
by: Vieira, Inês, et al.
Published: (2026)
by: Vieira, Inês, et al.
Published: (2026)
Multi-trait User Simulation with Adaptive Decoding for Conversational Task Assistants
by: Ferreira, Rafael, et al.
Published: (2024)
by: Ferreira, Rafael, et al.
Published: (2024)
VIGiA: Instructional Video Guidance via Dialogue Reasoning and Retrieval
by: Glória-Silva, Diogo, et al.
Published: (2026)
by: Glória-Silva, Diogo, et al.
Published: (2026)
MoralCLIP: Contrastive Alignment of Vision-and-Language Representations with Moral Foundations Theory
by: Condez, Ana Carolina, et al.
Published: (2025)
by: Condez, Ana Carolina, et al.
Published: (2025)
SRL-MAD: Structured Residual Latents for One-Class Morphing Attack Detection
by: Paulo, Diogo J., et al.
Published: (2026)
by: Paulo, Diogo J., et al.
Published: (2026)
FD-MAD: Frequency-Domain Residual Analysis for Face Morphing Attack Detection
by: Paulo, Diogo J., et al.
Published: (2026)
by: Paulo, Diogo J., et al.
Published: (2026)
GlórIA -- A Generative and Open Large Language Model for Portuguese
by: Lopes, Ricardo, et al.
Published: (2024)
by: Lopes, Ricardo, et al.
Published: (2024)
MAD-MAX: Modular And Diverse Malicious Attack MiXtures for Automated LLM Red Teaming
by: Schoepf, Stefan, et al.
Published: (2025)
by: Schoepf, Stefan, et al.
Published: (2025)
Adaptive Instruction Composition for Automated LLM Red-Teaming
by: Zymet, Jesse, et al.
Published: (2026)
by: Zymet, Jesse, et al.
Published: (2026)
AI in Insurance: Adaptive Questionnaires for Improved Risk Profiling
by: Silva, Diogo, et al.
Published: (2026)
by: Silva, Diogo, et al.
Published: (2026)
Automatic LLM Red Teaming
by: Belaire, Roman, et al.
Published: (2025)
by: Belaire, Roman, et al.
Published: (2025)
Red Teaming AI Red Teaming
by: Majumdar, Subhabrata, et al.
Published: (2025)
by: Majumdar, Subhabrata, et al.
Published: (2025)
AutoRedTeamer: Autonomous Red Teaming with Lifelong Attack Integration
by: Zhou, Andy, et al.
Published: (2025)
by: Zhou, Andy, et al.
Published: (2025)
Genesis: Evolving Attack Strategies for LLM Web Agent Red-Teaming
by: Zhang, Zheng, et al.
Published: (2025)
by: Zhang, Zheng, et al.
Published: (2025)
Red-Teaming LLM Multi-Agent Systems via Communication Attacks
by: He, Pengfei, et al.
Published: (2025)
by: He, Pengfei, et al.
Published: (2025)
Visual Exclusivity Attacks: Automatic Multimodal Red Teaming via Agentic Planning
by: Zhang, Yunbei, et al.
Published: (2026)
by: Zhang, Yunbei, et al.
Published: (2026)
AMALIA Technical Report: A Fully Open Source Large Language Model for European Portuguese
by: Simplício, Afonso, et al.
Published: (2026)
by: Simplício, Afonso, et al.
Published: (2026)
RedTopic: Toward Topic-Diverse Red Teaming of Large Language Models
by: Ding, Jiale, et al.
Published: (2025)
by: Ding, Jiale, et al.
Published: (2025)
Prompt Optimization and Evaluation for LLM Automated Red Teaming
by: Freenor, Michael, et al.
Published: (2025)
by: Freenor, Michael, et al.
Published: (2025)
ARMs: Adaptive Red-Teaming Agent against Multimodal Models with Plug-and-Play Attacks
by: Chen, Zhaorun, et al.
Published: (2025)
by: Chen, Zhaorun, et al.
Published: (2025)
StreetView-Waste: A Multi-Task Dataset for Urban Waste Management
by: Paulo, Diogo J., et al.
Published: (2025)
by: Paulo, Diogo J., et al.
Published: (2025)
SIRAJ: Diverse and Efficient Red-Teaming for LLM Agents via Distilled Structured Reasoning
by: Zhou, Kaiwen, et al.
Published: (2025)
by: Zhou, Kaiwen, et al.
Published: (2025)
TroubleLLM: Align to Red Team Expert
by: Xu, Zhuoer, et al.
Published: (2024)
by: Xu, Zhuoer, et al.
Published: (2024)
Autonomous Adversary: Red-Teaming in the age of LLM
by: Mamun, Mohammad, et al.
Published: (2026)
by: Mamun, Mohammad, et al.
Published: (2026)
pyMOE: Python library for designing and producing masks for Micro Optical Elements
by: Joao Cunha, et al.
Published: (2026)
by: Joao Cunha, et al.
Published: (2026)
pyMOE: Python library for designing and producing masks for Micro Optical Elements
by: Joao Cunha, et al.
Published: (2026)
by: Joao Cunha, et al.
Published: (2026)
Preheated inflation
by: Gorgulho, Diogo S., et al.
Published: (2025)
by: Gorgulho, Diogo S., et al.
Published: (2025)
Ruby Teaming: Improving Quality Diversity Search with Memory for Automated Red Teaming
by: Han, Vernon Toh Yan, et al.
Published: (2024)
by: Han, Vernon Toh Yan, et al.
Published: (2024)
MUZZLE: Adaptive Agentic Red-Teaming of Web Agents Against Indirect Prompt Injection Attacks
by: Syros, Georgios, et al.
Published: (2026)
by: Syros, Georgios, et al.
Published: (2026)
O lugar do camponês e questão agrária na Revolução Russa de 1917
by: Ana Claudia Diogo Tavares
Published: (2017)
by: Ana Claudia Diogo Tavares
Published: (2017)
Gastric Cancer Screening: Intention to Adhere and Patients' Perspective
by: João Carlos Silva, et al.
Published: (2024)
by: João Carlos Silva, et al.
Published: (2024)
Generating Coherent Sequences of Visual Illustrations for Real-World Manual Tasks
by: Bordalo, João, et al.
Published: (2024)
by: Bordalo, João, et al.
Published: (2024)
Flexibilização organizacional e empregabilidade individual: proposição de um modelo explicativo
by: Diogo Henrique Helal
Published: (2005)
by: Diogo Henrique Helal
Published: (2005)
ENTRE PRECEITOS FORDISTAS E DE FLEXIBILIZAÇÃO: UM ESTUDO DE CASO SOBRE O PROCESSO DE MUDANÇA DE PRODUÇÃO E TRABALHO NA CONSTRUÇÃO CIVIL
by: Diogo Henrique Helal
Published: (2007)
by: Diogo Henrique Helal
Published: (2007)
O discurso da empregabilidade: o que pensam a academia e o mundo empresarial
by: Diogo Henrique Helal
Published: (2011)
by: Diogo Henrique Helal
Published: (2011)
Superando a pobreza: o papel do capital social na Região Metropolitana de Belo Horizonte
by: Diogo Henrique Helal
Published: (2007)
by: Diogo Henrique Helal
Published: (2007)
Burocracia e inserção social: uma proposta para entender a gestão das organizações públicas no Brasil
by: Diogo Henrique Helal
Published: (2010)
by: Diogo Henrique Helal
Published: (2010)
Similar Items
-
TWIZ-v2: The Wizard of Multimodal Conversational-Stimulus
by: Ferreira, Rafael, et al.
Published: (2023) -
Plan-Grounded Large Language Models for Dual Goal Conversational Settings
by: Glória-Silva, Diogo, et al.
Published: (2024) -
Show and Guide: Instructional-Plan Grounded Vision and Language Model
by: Glória-Silva, Diogo, et al.
Published: (2024) -
ALBA: A European Portuguese Benchmark for Evaluating Language and Linguistic Dimensions in Generative LLMs
by: Vieira, Inês, et al.
Published: (2026) -
Multi-trait User Simulation with Adaptive Decoding for Conversational Task Assistants
by: Ferreira, Rafael, et al.
Published: (2024)