PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Du, Pengfei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Red Teaming Large Reasoning Models
von: Chen, Jiawei, et al.
Veröffentlicht: (2025)
von: Chen, Jiawei, et al.
Veröffentlicht: (2025)
CORVUS: Red-Teaming Hallucination Detectors via Internal Signal Camouflage in Large Language Models
von: Min, Nay Myat, et al.
Veröffentlicht: (2026)
von: Min, Nay Myat, et al.
Veröffentlicht: (2026)
BlackIce: A Containerized Red Teaming Toolkit for AI Security Testing
von: Kaplan, Caelin, et al.
Veröffentlicht: (2025)
von: Kaplan, Caelin, et al.
Veröffentlicht: (2025)
A Systematic Review of Algorithmic Red Teaming Methodologies for Assurance and Security of AI Applications
von: Srivastava, Shruti, et al.
Veröffentlicht: (2026)
von: Srivastava, Shruti, et al.
Veröffentlicht: (2026)
Red-Teaming Agent Execution Contexts: Open-World Security Evaluation on OpenClaw
von: Yao, Hongwei, et al.
Veröffentlicht: (2026)
von: Yao, Hongwei, et al.
Veröffentlicht: (2026)
When Search Goes Wrong: Red-Teaming Web-Augmented Large Language Models
von: Ou, Haoran, et al.
Veröffentlicht: (2025)
von: Ou, Haoran, et al.
Veröffentlicht: (2025)
Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment
von: Wang, Haoran, et al.
Veröffentlicht: (2023)
von: Wang, Haoran, et al.
Veröffentlicht: (2023)
Red-Teaming Coding Agents from a Tool-Invocation Perspective: An Empirical Security Assessment
von: Xie, Yuchong, et al.
Veröffentlicht: (2025)
von: Xie, Yuchong, et al.
Veröffentlicht: (2025)
Red-Teaming Claude Opus and ChatGPT-based Security Advisors for Trusted Execution Environments
von: Mukherjee, Kunal, et al.
Veröffentlicht: (2026)
von: Mukherjee, Kunal, et al.
Veröffentlicht: (2026)
Red Teaming AI Red Teaming
von: Majumdar, Subhabrata, et al.
Veröffentlicht: (2025)
von: Majumdar, Subhabrata, et al.
Veröffentlicht: (2025)
ClawTrap: A MITM-Based Red-Teaming Framework for Real-World OpenClaw Security Evaluation
von: Zhao, Haochen, et al.
Veröffentlicht: (2026)
von: Zhao, Haochen, et al.
Veröffentlicht: (2026)
Medical Multimodal Model Stealing Attacks via Adversarial Domain Alignment
von: Shen, Yaling, et al.
Veröffentlicht: (2025)
von: Shen, Yaling, et al.
Veröffentlicht: (2025)
AutoRedTeamer: Autonomous Red Teaming with Lifelong Attack Integration
von: Zhou, Andy, et al.
Veröffentlicht: (2025)
von: Zhou, Andy, et al.
Veröffentlicht: (2025)
RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent
von: Xu, Huiyu, et al.
Veröffentlicht: (2024)
von: Xu, Huiyu, et al.
Veröffentlicht: (2024)
Whispers of Wealth: Red-Teaming Google's Agent Payments Protocol via Prompt Injection
von: Debi, Tanusree, et al.
Veröffentlicht: (2026)
von: Debi, Tanusree, et al.
Veröffentlicht: (2026)
SAGE-RT: Synthetic Alignment data Generation for Safety Evaluation and Red Teaming
von: Kumar, Anurakt, et al.
Veröffentlicht: (2024)
von: Kumar, Anurakt, et al.
Veröffentlicht: (2024)
Automatic Red Teaming LLM-based Agents with Model Context Protocol Tools
von: He, Ping, et al.
Veröffentlicht: (2025)
von: He, Ping, et al.
Veröffentlicht: (2025)
A Red Teaming Roadmap Towards System-Level Safety
von: Wang, Zifan, et al.
Veröffentlicht: (2025)
von: Wang, Zifan, et al.
Veröffentlicht: (2025)
MonitoringBench: Semi-Automated Red-Teaming for Agent Monitoring
von: Jotautaitė, Monika, et al.
Veröffentlicht: (2026)
von: Jotautaitė, Monika, et al.
Veröffentlicht: (2026)
Co-Evolutionary Multi-Modal Alignment via Structured Adversarial Evolution
von: Shi, Guoxin, et al.
Veröffentlicht: (2026)
von: Shi, Guoxin, et al.
Veröffentlicht: (2026)
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models
von: Liu, Yanjiang, et al.
Veröffentlicht: (2025)
von: Liu, Yanjiang, et al.
Veröffentlicht: (2025)
AutoRISE: Agent-Driven Strategy Evolution for Red-Teaming Large Language Models
von: Gautam, Tanmay, et al.
Veröffentlicht: (2026)
von: Gautam, Tanmay, et al.
Veröffentlicht: (2026)
Redefining AI Red Teaming in the Agentic Era: From Weeks to Hours
von: Dheekonda, Raja Sekhar Rao, et al.
Veröffentlicht: (2026)
von: Dheekonda, Raja Sekhar Rao, et al.
Veröffentlicht: (2026)
Proteus: A Self-Evolving Red Team for Agent Skill Ecosystems
von: Zhou, Zhaojiacheng
Veröffentlicht: (2026)
von: Zhou, Zhaojiacheng
Veröffentlicht: (2026)
PyRIT: A Framework for Security Risk Identification and Red Teaming in Generative AI System
von: Munoz, Gary D. Lopez, et al.
Veröffentlicht: (2024)
von: Munoz, Gary D. Lopez, et al.
Veröffentlicht: (2024)
Incalmo: An Autonomous LLM-assisted System for Red Teaming Multi-Host Networks
von: Singer, Brian, et al.
Veröffentlicht: (2025)
von: Singer, Brian, et al.
Veröffentlicht: (2025)
Security-aware Semantic-driven ISAC via Paired Adversarial Residual Networks
von: Liu, Yu, et al.
Veröffentlicht: (2025)
von: Liu, Yu, et al.
Veröffentlicht: (2025)
Information Theoretic Adversarial Training of Large Language Models
von: Zhang, Yiwei, et al.
Veröffentlicht: (2026)
von: Zhang, Yiwei, et al.
Veröffentlicht: (2026)
From Firewalls to Frontiers: AI Red-Teaming is a Domain-Specific Evolution of Cyber Red-Teaming
von: Sinha, Anusha, et al.
Veröffentlicht: (2025)
von: Sinha, Anusha, et al.
Veröffentlicht: (2025)
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications
von: David, Isaac, et al.
Veröffentlicht: (2026)
von: David, Isaac, et al.
Veröffentlicht: (2026)
Trojan Horses in Recruiting: A Red-Teaming Case Study on Indirect Prompt Injection in Standard vs. Reasoning Models
von: Wirth, Manuel
Veröffentlicht: (2026)
von: Wirth, Manuel
Veröffentlicht: (2026)
Leveraging RAG for Training-Free Alignment of LLMs
von: Halloran, John T.
Veröffentlicht: (2026)
von: Halloran, John T.
Veröffentlicht: (2026)
(Security) Assertions by Large Language Models
von: Kande, Rahul, et al.
Veröffentlicht: (2023)
von: Kande, Rahul, et al.
Veröffentlicht: (2023)
MUZZLE: Adaptive Agentic Red-Teaming of Web Agents Against Indirect Prompt Injection Attacks
von: Syros, Georgios, et al.
Veröffentlicht: (2026)
von: Syros, Georgios, et al.
Veröffentlicht: (2026)
Secure Text Mail Encryption with Generative Adversarial Networks
von: Schelle, Alexej
Veröffentlicht: (2025)
von: Schelle, Alexej
Veröffentlicht: (2025)
Investigating Deep Watermark Security: An Adversarial Transferability Perspective
von: Qi, Biqing, et al.
Veröffentlicht: (2024)
von: Qi, Biqing, et al.
Veröffentlicht: (2024)
Large Empirical Case Study: Go-Explore adapted for AI Red Team Testing
von: Bhatt, Manish, et al.
Veröffentlicht: (2025)
von: Bhatt, Manish, et al.
Veröffentlicht: (2025)
Large Language Model-driven Security Assistant for Internet of Things via Chain-of-Thought
von: Zeng, Mingfei, et al.
Veröffentlicht: (2025)
von: Zeng, Mingfei, et al.
Veröffentlicht: (2025)
Evaluating Adversarial Vulnerabilities in Modern Large Language Models
von: Perel, Tom
Veröffentlicht: (2025)
von: Perel, Tom
Veröffentlicht: (2025)
Emerging Security Challenges of Large Language Models
von: Debar, Herve, et al.
Veröffentlicht: (2024)
von: Debar, Herve, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Red Teaming Large Reasoning Models
von: Chen, Jiawei, et al.
Veröffentlicht: (2025) -
CORVUS: Red-Teaming Hallucination Detectors via Internal Signal Camouflage in Large Language Models
von: Min, Nay Myat, et al.
Veröffentlicht: (2026) -
BlackIce: A Containerized Red Teaming Toolkit for AI Security Testing
von: Kaplan, Caelin, et al.
Veröffentlicht: (2025) -
A Systematic Review of Algorithmic Red Teaming Methodologies for Assurance and Security of AI Applications
von: Srivastava, Shruti, et al.
Veröffentlicht: (2026) -
Red-Teaming Agent Execution Contexts: Open-World Security Evaluation on OpenClaw
von: Yao, Hongwei, et al.
Veröffentlicht: (2026)