DecodingTrust-Agent Platform (DTap): A Controllable and Interactive Red-Teaming Platform for AI Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Zhaorun, Liu, Xun, Tong, Haibo, Guo, Chengquan, Nie, Yuzhou, Zhang, Jiawei, Kang, Mintong, Xu, Chejian, Liu, Qichang, Liu, Xiaogeng, Shi, Tianneng, Xiao, Chaowei, Koyejo, Sanmi, Liang, Percy, Guo, Wenbo, Song, Dawn, Li, Bo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models
by: Wang, Boxin, et al.
Published: (2023)
by: Wang, Boxin, et al.
Published: (2023)
BlueCodeAgent: A Blue Teaming Agent Enabled by Automated Red Teaming for CodeGen AI
by: Guo, Chengquan, et al.
Published: (2025)
by: Guo, Chengquan, et al.
Published: (2025)
ARMs: Adaptive Red-Teaming Agent against Multimodal Models with Plug-and-Play Attacks
by: Chen, Zhaorun, et al.
Published: (2025)
by: Chen, Zhaorun, et al.
Published: (2025)
Poly-Guard: Massive Multi-Domain Safety Policy-Grounded Guardrail Dataset
by: Kang, Mintong, et al.
Published: (2025)
by: Kang, Mintong, et al.
Published: (2025)
MetaAgent: Automatically Constructing Multi-Agent Systems Based on Finite State Machines
by: Zhang, Yaolun, et al.
Published: (2025)
by: Zhang, Yaolun, et al.
Published: (2025)
ShieldAgent: Shielding Agents via Verifiable Safety Policy Reasoning
by: Chen, Zhaorun, et al.
Published: (2025)
by: Chen, Zhaorun, et al.
Published: (2025)
VulnLLM-R: Specialized Reasoning LLM with Agent Scaffold for Vulnerability Detection
by: Nie, Yuzhou, et al.
Published: (2025)
by: Nie, Yuzhou, et al.
Published: (2025)
AgentVigil: Generic Black-Box Red-teaming for Indirect Prompt Injection against LLM Agents
by: Wang, Zhun, et al.
Published: (2025)
by: Wang, Zhun, et al.
Published: (2025)
EIA: Environmental Injection Attack on Generalist Web Agents for Privacy Leakage
by: Liao, Zeyi, et al.
Published: (2024)
by: Liao, Zeyi, et al.
Published: (2024)
AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases
by: Chen, Zhaorun, et al.
Published: (2024)
by: Chen, Zhaorun, et al.
Published: (2024)
RedCodeAgent: Automatic Red-teaming Agent against Diverse Code Agents
by: Guo, Chengquan, et al.
Published: (2025)
by: Guo, Chengquan, et al.
Published: (2025)
AutoDAN-Reasoning: Enhancing Strategies Exploration based Jailbreak Attacks with Test-Time Scaling
by: Liu, Xiaogeng, et al.
Published: (2025)
by: Liu, Xiaogeng, et al.
Published: (2025)
AdvAgent: Controllable Blackbox Red-teaming on Web Agents
by: Xu, Chejian, et al.
Published: (2024)
by: Xu, Chejian, et al.
Published: (2024)
Progent: Securing AI Agents with Privilege Control
by: Shi, Tianneng, et al.
Published: (2025)
by: Shi, Tianneng, et al.
Published: (2025)
AdvWave: Stealthy Adversarial Jailbreak Attack against Large Audio-Language Models
by: Kang, Mintong, et al.
Published: (2024)
by: Kang, Mintong, et al.
Published: (2024)
RedCode: Risky Code Execution and Generation Benchmark for Code Agents
by: Guo, Chengquan, et al.
Published: (2024)
by: Guo, Chengquan, et al.
Published: (2024)
ROM: Real-time Overthinking Mitigation via Streaming Detection and Intervention
by: Wang, Xinyan, et al.
Published: (2026)
by: Wang, Xinyan, et al.
Published: (2026)
OET: Optimization-based prompt injection Evaluation Toolkit
by: Pan, Jinsheng, et al.
Published: (2025)
by: Pan, Jinsheng, et al.
Published: (2025)
RePD: Defending Jailbreak Attack through a Retrieval-based Prompt Decomposition Process
by: Wang, Peiran, et al.
Published: (2024)
by: Wang, Peiran, et al.
Published: (2024)
LeakAgent: RL-based Red-teaming Agent for LLM Privacy Leakage
by: Nie, Yuzhou, et al.
Published: (2024)
by: Nie, Yuzhou, et al.
Published: (2024)
JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks
by: Luo, Weidi, et al.
Published: (2024)
by: Luo, Weidi, et al.
Published: (2024)
DRIFT: Dynamic Rule-Based Defense with Injection Isolation for Securing LLM Agents
by: Li, Hao, et al.
Published: (2025)
by: Li, Hao, et al.
Published: (2025)
AGrail: A Lifelong Agent Guardrail with Effective and Adaptive Safety Detection
by: Luo, Weidi, et al.
Published: (2025)
by: Luo, Weidi, et al.
Published: (2025)
Extracting books from production language models
by: Ahmed, Ahmed, et al.
Published: (2026)
by: Ahmed, Ahmed, et al.
Published: (2026)
AutoRedTeamer: Autonomous Red Teaming with Lifelong Attack Integration
by: Zhou, Andy, et al.
Published: (2025)
by: Zhou, Andy, et al.
Published: (2025)
SafeGen-Bench: Benchmarking Safety in Image-Conditioned Text-to-Video Generation
by: Ma, Yingzi, et al.
Published: (2026)
by: Ma, Yingzi, et al.
Published: (2026)
AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
by: Liu, Xiaogeng, et al.
Published: (2023)
by: Liu, Xiaogeng, et al.
Published: (2023)
SeCodePLT: A Unified Platform for Evaluating the Security of Code GenAI
by: Nie, Yuzhou, et al.
Published: (2024)
by: Nie, Yuzhou, et al.
Published: (2024)
Causally Inspired Regularization Enables Domain General Representations
by: Salaudeen, Olawale, et al.
Published: (2024)
by: Salaudeen, Olawale, et al.
Published: (2024)
Let's Measure Information Step-by-Step: AI-Based Evaluation Beyond Vibes
by: Robertson, Zachary, et al.
Published: (2025)
by: Robertson, Zachary, et al.
Published: (2025)
CURE: Cultural Understanding and Reasoning Evaluation - A Framework for "Thick" Culture Alignment Evaluation in LLMs
by: Vo, Truong, et al.
Published: (2025)
by: Vo, Truong, et al.
Published: (2025)
LEGENT: Open Platform for Embodied Agents
by: Cheng, Zhili, et al.
Published: (2024)
by: Cheng, Zhili, et al.
Published: (2024)
Reliable and Efficient Amortized Model-based Evaluation
by: Truong, Sang, et al.
Published: (2025)
by: Truong, Sang, et al.
Published: (2025)
SpecEval: Evaluating Model Adherence to Behavior Specifications
by: Ahmed, Ahmed, et al.
Published: (2025)
by: Ahmed, Ahmed, et al.
Published: (2025)
The Optimization Paradox in Clinical AI Multi-Agent Systems
by: Bedi, Suhana, et al.
Published: (2025)
by: Bedi, Suhana, et al.
Published: (2025)
Welfare, Improvability, and Variance: A Principal-Agent Approach to Optimal Benchmark Item Aggregation
by: Haupt, Andreas, et al.
Published: (2026)
by: Haupt, Andreas, et al.
Published: (2026)
TeamTR: Trust-Region Fine-Tuning for Multi-Agent LLM Coordination
by: Xie, Yi, et al.
Published: (2026)
by: Xie, Yi, et al.
Published: (2026)
Hybrid Team Tetris: A New Platform For Hybrid Multi-Agent, Multi-Human Teaming
by: Mcdowell, Kaleb, et al.
Published: (2025)
by: Mcdowell, Kaleb, et al.
Published: (2025)
CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale
by: Wang, Zhun, et al.
Published: (2025)
by: Wang, Zhun, et al.
Published: (2025)
The Next Paradigm Is User-Centric Agent, Not Platform-Centric Service
by: Zhang, Luankang, et al.
Published: (2026)
by: Zhang, Luankang, et al.
Published: (2026)
Similar Items
-
DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models
by: Wang, Boxin, et al.
Published: (2023) -
BlueCodeAgent: A Blue Teaming Agent Enabled by Automated Red Teaming for CodeGen AI
by: Guo, Chengquan, et al.
Published: (2025) -
ARMs: Adaptive Red-Teaming Agent against Multimodal Models with Plug-and-Play Attacks
by: Chen, Zhaorun, et al.
Published: (2025) -
Poly-Guard: Massive Multi-Domain Safety Policy-Grounded Guardrail Dataset
by: Kang, Mintong, et al.
Published: (2025) -
MetaAgent: Automatically Constructing Multi-Agent Systems Based on Finite State Machines
by: Zhang, Yaolun, et al.
Published: (2025)