Learning-Based Automated Adversarial Red-Teaming for Robustness Evaluation of Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Wei, Zhang, Chen, Hanxuan, Hu, Peilu, Wei, Zhenyuan, Liang, Chenwei, Luo, Jing, Ni, Ziyi, Yan, Hao, Mei, Li, Lang, Shengning, Lu, Kuan, Xiao, Xi, Han, Zhimo, Wang, Yijin, Zhang, Yichao, Yang, Chen, Hao, Junfeng, Gu, Jiayi, Bao, Riyang, Wang, Mu-Jiang-Shan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AOI: Context-Aware Multi-Agent Operations via Dynamic Scheduling and Hierarchical Memory Compression
by: Bai, Zishan, et al.
Published: (2025)
by: Bai, Zishan, et al.
Published: (2025)
On the Elementary Proof of the Inverse Erdős-Heilbronn Problem
by: Zhang, Shengning
Published: (2024)
by: Zhang, Shengning
Published: (2024)
Effects of ABO‐Incompatible Blood Transfusion on Immune Response and Rejection After Organ Transplantation
by: Peilu Hu, et al.
Published: (2026)
by: Peilu Hu, et al.
Published: (2026)
Illuminating the genomic frontier of invasive non‐typhoidal Salmonella infections
by: Hao Wang, et al.
Published: (2025)
by: Hao Wang, et al.
Published: (2025)
AutoRed: A Free-form Adversarial Prompt Generation Framework for Automated Red Teaming
by: Diao, Muxi, et al.
Published: (2025)
by: Diao, Muxi, et al.
Published: (2025)
AutoPrompt: Automated Red-Teaming of Text-to-Image Models via LLM-Driven Adversarial Prompts
by: Liu, Yufan, et al.
Published: (2025)
by: Liu, Yufan, et al.
Published: (2025)
ShieldLearner: A New Paradigm for Jailbreak Attack Defense in LLMs
by: Ni, Ziyi, et al.
Published: (2025)
by: Ni, Ziyi, et al.
Published: (2025)
SafeSearch: Automated Red-Teaming of LLM-Based Search Agents
by: Dong, Jianshuo, et al.
Published: (2025)
by: Dong, Jianshuo, et al.
Published: (2025)
Embodied Red Teaming for Auditing Robotic Foundation Models
by: Karnik, Sathwik, et al.
Published: (2024)
by: Karnik, Sathwik, et al.
Published: (2024)
Cover Feature: Efficient Hydrodeoxygenation of Lignin‐Derived Phenolic Compounds Over Ru‐Based Catalyst with Biochar and Al2O3 as Composite Support (ChemSusChem 1/2025)
by: Tao Yin, et al.
Published: (2025)
by: Tao Yin, et al.
Published: (2025)
AgenticRed: Evolving Agentic Systems for Red-Teaming
by: Yuan, Jiayi, et al.
Published: (2026)
by: Yuan, Jiayi, et al.
Published: (2026)
Track A*: Fast Visibility-Aware Trajectory Planning for Active Target Tracking
by: Chen, Hanxuan, et al.
Published: (2026)
by: Chen, Hanxuan, et al.
Published: (2026)
Genesis: Evolving Attack Strategies for LLM Web Agent Red-Teaming
by: Zhang, Zheng, et al.
Published: (2025)
by: Zhang, Zheng, et al.
Published: (2025)
Autonomous Adversary: Red-Teaming in the age of LLM
by: Mamun, Mohammad, et al.
Published: (2026)
by: Mamun, Mohammad, et al.
Published: (2026)
Unintended Negative Impacts of Promotional Language in Patent Evaluation
by: Zhao, Bingkun, et al.
Published: (2026)
by: Zhao, Bingkun, et al.
Published: (2026)
A Supramolecular System of Bioactive Ionic Liquid, Peony Extract, and Peptide for Enhanced Permeability and Synergistic Skincare Benefits
by: Mi Wang, et al.
Published: (2025)
by: Mi Wang, et al.
Published: (2025)
Step-by-Step Causality: Transparent Causal Discovery with Multi-Agent Tree-Query and Adversarial Confidence Estimation
by: Ding, Ziyi, et al.
Published: (2026)
by: Ding, Ziyi, et al.
Published: (2026)
Jailbreaking and Mitigation of Vulnerabilities in Large Language Models
by: Peng, Benji, et al.
Published: (2024)
by: Peng, Benji, et al.
Published: (2024)
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming
by: Zheng, Xiang, et al.
Published: (2025)
by: Zheng, Xiang, et al.
Published: (2025)
DRL-Based Beam Positioning for LEO Satellite Constellations with Weighted Least Squares
by: Chou, Po-Heng, et al.
Published: (2025)
by: Chou, Po-Heng, et al.
Published: (2025)
Against The Achilles' Heel: A Survey on Red Teaming for Generative Models
by: Lin, Lizhi, et al.
Published: (2024)
by: Lin, Lizhi, et al.
Published: (2024)
The Geographical Differences in the Morphology of Diptychus maculatus: Environmental Driving Factors and Adaptive Evolution
by: Yichao Hao, et al.
Published: (2026)
by: Yichao Hao, et al.
Published: (2026)
Red Teaming AI Red Teaming
by: Majumdar, Subhabrata, et al.
Published: (2025)
by: Majumdar, Subhabrata, et al.
Published: (2025)
Learning to Learn with Quantum Optimization via Quantum Neural Networks
by: Chen, Kuan-Cheng, et al.
Published: (2025)
by: Chen, Kuan-Cheng, et al.
Published: (2025)
TroubleLLM: Align to Red Team Expert
by: Xu, Zhuoer, et al.
Published: (2024)
by: Xu, Zhuoer, et al.
Published: (2024)
Realizing Video Summarization from the Path of Language-based Semantic Understanding
by: Mu, Kuan-Chen, et al.
Published: (2024)
by: Mu, Kuan-Chen, et al.
Published: (2024)
Preparation and Study of Highly Hydrophobic and Strongly Oleophilic Reticulated Polyurethane Foam Surface Modified With Polydopamine and Polyhedral Oligomeric Silsesquioxane
by: Longyu Hao, et al.
Published: (2025)
by: Longyu Hao, et al.
Published: (2025)
Quadratic-form Optimal Transport
by: Wang, Ruodu, et al.
Published: (2025)
by: Wang, Ruodu, et al.
Published: (2025)
Simultaneous Optimal Transport
by: Wang, Ruodu, et al.
Published: (2022)
by: Wang, Ruodu, et al.
Published: (2022)
Considering the Difference in Utility Functions of Team Players in Adversarial Team Games
by: Zhang, Youzhi
Published: (2025)
by: Zhang, Youzhi
Published: (2025)
Persona-Conditioned Adversarial Prompting: Multi-Identity Red-Teaming for Adversarial Discovery and Mitigation
by: Morasso, Cristian, et al.
Published: (2026)
by: Morasso, Cristian, et al.
Published: (2026)
RSCC: A Large-Scale Remote Sensing Change Caption Dataset for Disaster Events
by: Chen, Zhenyuan, et al.
Published: (2025)
by: Chen, Zhenyuan, et al.
Published: (2025)
Rethinking Inductive Bias in Geographically Neural Network Weighted Regression
by: Chen, Zhenyuan
Published: (2025)
by: Chen, Zhenyuan
Published: (2025)
Towards Dynamic Resource Allocation and Client Scheduling in Hierarchical Federated Learning: A Two-Phase Deep Reinforcement Learning Approach
by: Chen, Xiaojing, et al.
Published: (2024)
by: Chen, Xiaojing, et al.
Published: (2024)
ReToMe-VA: Recursive Token Merging for Video Diffusion-based Unrestricted Adversarial Attack
by: Gao, Ziyi, et al.
Published: (2024)
by: Gao, Ziyi, et al.
Published: (2024)
Metabolomics and Microbiomics Perspectives Reveal the Regulatory Pathways of Monaphilone B Derived From Red Yeast Rice on Alcoholic Liver Injury in Mice
by: Li Wu, et al.
Published: (2025)
by: Li Wu, et al.
Published: (2025)
RedTopic: Toward Topic-Diverse Red Teaming of Large Language Models
by: Ding, Jiale, et al.
Published: (2025)
by: Ding, Jiale, et al.
Published: (2025)
FedDTG:Federated Data-Free Knowledge Distillation via Three-Player Generative Adversarial Networks
by: Gao, Lingzhi, et al.
Published: (2022)
by: Gao, Lingzhi, et al.
Published: (2022)
Online Covariance Estimation in Averaged SGD: Improved Batch-Mean Rates and Minimax Optimality via Trajectory Regression
by: Ni, Yijin, et al.
Published: (2026)
by: Ni, Yijin, et al.
Published: (2026)
A Uniform Concentration Inequality for Kernel-Based Two-Sample Statistics
by: Ni, Yijin, et al.
Published: (2024)
by: Ni, Yijin, et al.
Published: (2024)
Similar Items
-
AOI: Context-Aware Multi-Agent Operations via Dynamic Scheduling and Hierarchical Memory Compression
by: Bai, Zishan, et al.
Published: (2025) -
On the Elementary Proof of the Inverse Erdős-Heilbronn Problem
by: Zhang, Shengning
Published: (2024) -
Effects of ABO‐Incompatible Blood Transfusion on Immune Response and Rejection After Organ Transplantation
by: Peilu Hu, et al.
Published: (2026) -
Illuminating the genomic frontier of invasive non‐typhoidal Salmonella infections
by: Hao Wang, et al.
Published: (2025) -
AutoRed: A Free-form Adversarial Prompt Generation Framework for Automated Red Teaming
by: Diao, Muxi, et al.
Published: (2025)