RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xu, Huiyu, Zhang, Wenhui, Wang, Zhibo, Xiao, Feng, Zheng, Rui, Feng, Yunhe, Ba, Zhongjie, Ren, Kui |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Resource Consumption Red-Teaming for Large Vision-Language Models
von: Gao, Haoran, et al.
Veröffentlicht: (2025)
von: Gao, Haoran, et al.
Veröffentlicht: (2025)
T-MAP: Red-Teaming LLM Agents with Trajectory-aware Evolutionary Search
von: Lee, Hyomin, et al.
Veröffentlicht: (2026)
von: Lee, Hyomin, et al.
Veröffentlicht: (2026)
GenBreak: Red Teaming Text-to-Image Generators Using Large Language Models
von: Wang, Zilong, et al.
Veröffentlicht: (2025)
von: Wang, Zilong, et al.
Veröffentlicht: (2025)
Learning-Based Automated Adversarial Red-Teaming for Robustness Evaluation of Large Language Models
von: Wei, Zhang, et al.
Veröffentlicht: (2025)
von: Wei, Zhang, et al.
Veröffentlicht: (2025)
Towards Label-Only Membership Inference Attack against Pre-trained Large Language Models
von: He, Yu, et al.
Veröffentlicht: (2025)
von: He, Yu, et al.
Veröffentlicht: (2025)
SafeSearch: Automated Red-Teaming of LLM-Based Search Agents
von: Dong, Jianshuo, et al.
Veröffentlicht: (2025)
von: Dong, Jianshuo, et al.
Veröffentlicht: (2025)
LoopTrap: Termination Poisoning Attacks on LLM Agents
von: Xu, Huiyu, et al.
Veröffentlicht: (2026)
von: Xu, Huiyu, et al.
Veröffentlicht: (2026)
Effective Red-Teaming of Policy-Adherent Agents
von: Nakash, Itay, et al.
Veröffentlicht: (2025)
von: Nakash, Itay, et al.
Veröffentlicht: (2025)
AdvAgent: Controllable Blackbox Red-teaming on Web Agents
von: Xu, Chejian, et al.
Veröffentlicht: (2024)
von: Xu, Chejian, et al.
Veröffentlicht: (2024)
Can Small Language Models Reliably Resist Jailbreak Attacks? A Comprehensive Evaluation
von: Zhang, Wenhui, et al.
Veröffentlicht: (2025)
von: Zhang, Wenhui, et al.
Veröffentlicht: (2025)
Automated Progressive Red Teaming
von: Jiang, Bojian, et al.
Veröffentlicht: (2024)
von: Jiang, Bojian, et al.
Veröffentlicht: (2024)
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models
von: Liu, Yanjiang, et al.
Veröffentlicht: (2025)
von: Liu, Yanjiang, et al.
Veröffentlicht: (2025)
Jailbreak-Zero: A Path to Pareto Optimal Red Teaming for Large Language Models
von: Hu, Kai, et al.
Veröffentlicht: (2025)
von: Hu, Kai, et al.
Veröffentlicht: (2025)
RedTWIZ: Diverse LLM Red Teaming via Adaptive Attack Planning
von: Horal, Artur, et al.
Veröffentlicht: (2025)
von: Horal, Artur, et al.
Veröffentlicht: (2025)
Operationalizing a Threat Model for Red-Teaming Large Language Models (LLMs)
von: Verma, Apurv, et al.
Veröffentlicht: (2024)
von: Verma, Apurv, et al.
Veröffentlicht: (2024)
SIRAJ: Diverse and Efficient Red-Teaming for LLM Agents via Distilled Structured Reasoning
von: Zhou, Kaiwen, et al.
Veröffentlicht: (2025)
von: Zhou, Kaiwen, et al.
Veröffentlicht: (2025)
Interpretable LLM Guardrails via Sparse Representation Steering
von: He, Zeqing, et al.
Veröffentlicht: (2025)
von: He, Zeqing, et al.
Veröffentlicht: (2025)
DiveR-CT: Diversity-enhanced Red Teaming Large Language Model Assistants with Relaxing Constraints
von: Zhao, Andrew, et al.
Veröffentlicht: (2024)
von: Zhao, Andrew, et al.
Veröffentlicht: (2024)
Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment
von: Wang, Haoran, et al.
Veröffentlicht: (2023)
von: Wang, Haoran, et al.
Veröffentlicht: (2023)
Prompt Optimization and Evaluation for LLM Automated Red Teaming
von: Freenor, Michael, et al.
Veröffentlicht: (2025)
von: Freenor, Michael, et al.
Veröffentlicht: (2025)
Red-Teaming LLM Multi-Agent Systems via Communication Attacks
von: He, Pengfei, et al.
Veröffentlicht: (2025)
von: He, Pengfei, et al.
Veröffentlicht: (2025)
Training a General Purpose Automated Red Teaming Model
von: Padmakumar, Aishwarya, et al.
Veröffentlicht: (2026)
von: Padmakumar, Aishwarya, et al.
Veröffentlicht: (2026)
Autonomous Adversary: Red-Teaming in the age of LLM
von: Mamun, Mohammad, et al.
Veröffentlicht: (2026)
von: Mamun, Mohammad, et al.
Veröffentlicht: (2026)
AutoRISE: Agent-Driven Strategy Evolution for Red-Teaming Large Language Models
von: Gautam, Tanmay, et al.
Veröffentlicht: (2026)
von: Gautam, Tanmay, et al.
Veröffentlicht: (2026)
RerouteGuard: Understanding and Mitigating Adversarial Risks for LLM Routing
von: Zhang, Wenhui, et al.
Veröffentlicht: (2026)
von: Zhang, Wenhui, et al.
Veröffentlicht: (2026)
ContextualJailbreak: Evolutionary Red-Teaming via Simulated Conversational Priming
von: Béjar, Mario Rodríguez, et al.
Veröffentlicht: (2026)
von: Béjar, Mario Rodríguez, et al.
Veröffentlicht: (2026)
Holistic Automated Red Teaming for Large Language Models through Top-Down Test Case Generation and Multi-turn Interaction
von: Zhang, Jinchuan, et al.
Veröffentlicht: (2024)
von: Zhang, Jinchuan, et al.
Veröffentlicht: (2024)
Text-CRS: A Generalized Certified Robustness Framework against Textual Adversarial Attacks
von: Zhang, Xinyu, et al.
Veröffentlicht: (2023)
von: Zhang, Xinyu, et al.
Veröffentlicht: (2023)
MIRAGE: Multimodal Immersive Reasoning and Guided Exploration for Red-Team Jailbreak Attacks
von: You, Wenhao, et al.
Veröffentlicht: (2025)
von: You, Wenhao, et al.
Veröffentlicht: (2025)
Red-Teaming Agent Execution Contexts: Open-World Security Evaluation on OpenClaw
von: Yao, Hongwei, et al.
Veröffentlicht: (2026)
von: Yao, Hongwei, et al.
Veröffentlicht: (2026)
SkillAttack: Automated Red Teaming of Agent Skills through Attack Path Refinement
von: Duan, Zenghao, et al.
Veröffentlicht: (2026)
von: Duan, Zenghao, et al.
Veröffentlicht: (2026)
PT-Mark: Invisible Watermarking for Text-to-image Diffusion Models via Semantic-aware Pivotal Tuning
von: Wang, Yaopeng, et al.
Veröffentlicht: (2025)
von: Wang, Yaopeng, et al.
Veröffentlicht: (2025)
AutoRedTeamer: Autonomous Red Teaming with Lifelong Attack Integration
von: Zhou, Andy, et al.
Veröffentlicht: (2025)
von: Zhou, Andy, et al.
Veröffentlicht: (2025)
Atoxia: Red-teaming Large Language Models with Target Toxic Answers
von: Du, Yuhao, et al.
Veröffentlicht: (2024)
von: Du, Yuhao, et al.
Veröffentlicht: (2024)
InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents
von: Zhan, Qiusi, et al.
Veröffentlicht: (2024)
von: Zhan, Qiusi, et al.
Veröffentlicht: (2024)
Red Teaming the Mind of the Machine: A Systematic Evaluation of Prompt Injection and Jailbreak Vulnerabilities in LLMs
von: Pathade, Chetan
Veröffentlicht: (2025)
von: Pathade, Chetan
Veröffentlicht: (2025)
Automatic Red Teaming LLM-based Agents with Model Context Protocol Tools
von: He, Ping, et al.
Veröffentlicht: (2025)
von: He, Ping, et al.
Veröffentlicht: (2025)
Multi-Agent Collaboration in Incident Response with Large Language Models
von: Liu, Zefang
Veröffentlicht: (2024)
von: Liu, Zefang
Veröffentlicht: (2024)
MonitoringBench: Semi-Automated Red-Teaming for Agent Monitoring
von: Jotautaitė, Monika, et al.
Veröffentlicht: (2026)
von: Jotautaitė, Monika, et al.
Veröffentlicht: (2026)
A Certified Robust Watermark For Large Language Models
von: Feng, Xianheng, et al.
Veröffentlicht: (2024)
von: Feng, Xianheng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Resource Consumption Red-Teaming for Large Vision-Language Models
von: Gao, Haoran, et al.
Veröffentlicht: (2025) -
T-MAP: Red-Teaming LLM Agents with Trajectory-aware Evolutionary Search
von: Lee, Hyomin, et al.
Veröffentlicht: (2026) -
GenBreak: Red Teaming Text-to-Image Generators Using Large Language Models
von: Wang, Zilong, et al.
Veröffentlicht: (2025) -
Learning-Based Automated Adversarial Red-Teaming for Robustness Evaluation of Large Language Models
von: Wei, Zhang, et al.
Veröffentlicht: (2025) -
Towards Label-Only Membership Inference Attack against Pre-trained Large Language Models
von: He, Yu, et al.
Veröffentlicht: (2025)