DREAM: Dynamic Red-teaming across Environments for AI Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lu, Liming, Gu, Xiang, Huang, Junyu, Du, Jiawei, Zheng, Xu, Liu, Yunhuai, Zhou, Yongbin, Pang, Shuchao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SafeSteer: Adaptive Subspace Steering for Efficient Jailbreak Defense in Vision-Language Models
von: Zeng, Xiyu, et al.
Veröffentlicht: (2025)
von: Zeng, Xiyu, et al.
Veröffentlicht: (2025)
Reconstruction of Differentially Private Text Sanitization via Large Language Models
von: Pang, Shuchao, et al.
Veröffentlicht: (2024)
von: Pang, Shuchao, et al.
Veröffentlicht: (2024)
AgentRAE: Remote Action Execution through Notification-based Visual Backdoors against Screenshots-based Mobile GUI Agents
von: Luo, Yutao, et al.
Veröffentlicht: (2026)
von: Luo, Yutao, et al.
Veröffentlicht: (2026)
STARE: Step-wise Temporal Alignment and Red-teaming Engine for Multi-modal Toxicity Attack
von: Mao, Xutao, et al.
Veröffentlicht: (2026)
von: Mao, Xutao, et al.
Veröffentlicht: (2026)
Atoxia: Red-teaming Large Language Models with Target Toxic Answers
von: Du, Yuhao, et al.
Veröffentlicht: (2024)
von: Du, Yuhao, et al.
Veröffentlicht: (2024)
AdvAgent: Controllable Blackbox Red-teaming on Web Agents
von: Xu, Chejian, et al.
Veröffentlicht: (2024)
von: Xu, Chejian, et al.
Veröffentlicht: (2024)
AJAR: Adaptive Jailbreak Architecture for Red-teaming
von: Dou, Yipu, et al.
Veröffentlicht: (2026)
von: Dou, Yipu, et al.
Veröffentlicht: (2026)
CIARD: Cyclic Iterative Adversarial Robustness Distillation
von: Lu, Liming, et al.
Veröffentlicht: (2025)
von: Lu, Liming, et al.
Veröffentlicht: (2025)
HoneyWin: High-Interaction Windows Honeypot in Enterprise Environment
von: Aung, Yan Lin, et al.
Veröffentlicht: (2025)
von: Aung, Yan Lin, et al.
Veröffentlicht: (2025)
ASTRA: Autonomous Spatial-Temporal Red-teaming for AI Software Assistants
von: Xu, Xiangzhe, et al.
Veröffentlicht: (2025)
von: Xu, Xiangzhe, et al.
Veröffentlicht: (2025)
ART: Automatic Red-teaming for Text-to-Image Models to Protect Benign Users
von: Li, Guanlin, et al.
Veröffentlicht: (2024)
von: Li, Guanlin, et al.
Veröffentlicht: (2024)
DREAM: Scalable Red Teaming for Text-to-Image Generative Systems via Distribution Modeling
von: Li, Boheng, et al.
Veröffentlicht: (2025)
von: Li, Boheng, et al.
Veröffentlicht: (2025)
Detecting Scams Using Large Language Models
von: Jiang, Liming
Veröffentlicht: (2024)
von: Jiang, Liming
Veröffentlicht: (2024)
Utilizing Large LanguageModels to Detect Privacy Leaks in Mini-App Code
von: Jiang, Liming
Veröffentlicht: (2024)
von: Jiang, Liming
Veröffentlicht: (2024)
AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases
von: Chen, Zhaorun, et al.
Veröffentlicht: (2024)
von: Chen, Zhaorun, et al.
Veröffentlicht: (2024)
Red-teaming the Multimodal Reasoning: Jailbreaking Vision-Language Models via Cross-modal Entanglement Attacks
von: Yan, Yu, et al.
Veröffentlicht: (2026)
von: Yan, Yu, et al.
Veröffentlicht: (2026)
A Reward-driven Automated Webshell Malicious-code Generator for Red-teaming
von: Ding, Yizhong
Veröffentlicht: (2025)
von: Ding, Yizhong
Veröffentlicht: (2025)
MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teaming
von: Guo, Weiyang, et al.
Veröffentlicht: (2025)
von: Guo, Weiyang, et al.
Veröffentlicht: (2025)
Mastering AI: Big Data, Deep Learning, and the Evolution of Large Language Models -- Blockchain and Applications
von: Feng, Pohsun, et al.
Veröffentlicht: (2024)
von: Feng, Pohsun, et al.
Veröffentlicht: (2024)
Practically implementing an LLM-supported collaborative vulnerability remediation process: a team-based approach
von: Wang, Xiaoqing, et al.
Veröffentlicht: (2024)
von: Wang, Xiaoqing, et al.
Veröffentlicht: (2024)
Can LLMs Handle WebShell Detection? Overcoming Detection Challenges with Behavioral Function-Aware Framework
von: Han, Feijiang, et al.
Veröffentlicht: (2025)
von: Han, Feijiang, et al.
Veröffentlicht: (2025)
Red Teaming Methodology for Design Obfuscation
von: Liu, Yuntao, et al.
Veröffentlicht: (2025)
von: Liu, Yuntao, et al.
Veröffentlicht: (2025)
AgentVigil: Generic Black-Box Red-teaming for Indirect Prompt Injection against LLM Agents
von: Wang, Zhun, et al.
Veröffentlicht: (2025)
von: Wang, Zhun, et al.
Veröffentlicht: (2025)
SkillAttack: Automated Red Teaming of Agent Skills through Attack Path Refinement
von: Duan, Zenghao, et al.
Veröffentlicht: (2026)
von: Duan, Zenghao, et al.
Veröffentlicht: (2026)
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments
von: Zhang, Chiyu, et al.
Veröffentlicht: (2026)
von: Zhang, Chiyu, et al.
Veröffentlicht: (2026)
Red Teaming Large Reasoning Models
von: Chen, Jiawei, et al.
Veröffentlicht: (2025)
von: Chen, Jiawei, et al.
Veröffentlicht: (2025)
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training
von: Du, Pengfei
Veröffentlicht: (2025)
von: Du, Pengfei
Veröffentlicht: (2025)
From Thinker to Society: Security in Hierarchical Autonomy Evolution of AI Agents
von: Zhang, Xiaolei, et al.
Veröffentlicht: (2026)
von: Zhang, Xiaolei, et al.
Veröffentlicht: (2026)
Governing Dynamic Capabilities: Cryptographic Binding and Reproducibility Verification for AI Agent Tool Use
von: Zhou, Ziling
Veröffentlicht: (2026)
von: Zhou, Ziling
Veröffentlicht: (2026)
Demo: ViolentUTF as An Accessible Platform for Generative AI Red Teaming
von: Nguyen, Tam n.
Veröffentlicht: (2025)
von: Nguyen, Tam n.
Veröffentlicht: (2025)
RedShell: A Generative AI-Based Approach to Ethical Hacking
von: Bessa, Ricardo, et al.
Veröffentlicht: (2026)
von: Bessa, Ricardo, et al.
Veröffentlicht: (2026)
GenBreak: Red Teaming Text-to-Image Generators Using Large Language Models
von: Wang, Zilong, et al.
Veröffentlicht: (2025)
von: Wang, Zilong, et al.
Veröffentlicht: (2025)
LeakAgent: RL-based Red-teaming Agent for LLM Privacy Leakage
von: Nie, Yuzhou, et al.
Veröffentlicht: (2024)
von: Nie, Yuzhou, et al.
Veröffentlicht: (2024)
PA-CFL: Privacy-Adaptive Clustered Federated Learning for Transformer-Based Sales Forecasting on Heterogeneous Retail Data
von: Long, Yunbo, et al.
Veröffentlicht: (2025)
von: Long, Yunbo, et al.
Veröffentlicht: (2025)
Approximate Gaussian Mapping for Generative Image Steganography
von: Xu, Yuhua, et al.
Veröffentlicht: (2025)
von: Xu, Yuhua, et al.
Veröffentlicht: (2025)
Jailbreaking Commercial Black-Box LLMs with Explicitly Harmful Prompts
von: Zhang, Chiyu, et al.
Veröffentlicht: (2025)
von: Zhang, Chiyu, et al.
Veröffentlicht: (2025)
Multimodal Robust Prompt Distillation for 3D Point Cloud Models
von: Gu, Xiang, et al.
Veröffentlicht: (2025)
von: Gu, Xiang, et al.
Veröffentlicht: (2025)
UDora: A Unified Red Teaming Framework against LLM Agents by Dynamically Hijacking Their Own Reasoning
von: Zhang, Jiawei, et al.
Veröffentlicht: (2025)
von: Zhang, Jiawei, et al.
Veröffentlicht: (2025)
ModelShield: Adaptive and Robust Watermark against Model Extraction Attack
von: Pang, Kaiyi, et al.
Veröffentlicht: (2024)
von: Pang, Kaiyi, et al.
Veröffentlicht: (2024)
Shifting-Merging: Secure, High-Capacity and Efficient Steganography via Large Language Models
von: Bai, Minhao, et al.
Veröffentlicht: (2025)
von: Bai, Minhao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SafeSteer: Adaptive Subspace Steering for Efficient Jailbreak Defense in Vision-Language Models
von: Zeng, Xiyu, et al.
Veröffentlicht: (2025) -
Reconstruction of Differentially Private Text Sanitization via Large Language Models
von: Pang, Shuchao, et al.
Veröffentlicht: (2024) -
AgentRAE: Remote Action Execution through Notification-based Visual Backdoors against Screenshots-based Mobile GUI Agents
von: Luo, Yutao, et al.
Veröffentlicht: (2026) -
STARE: Step-wise Temporal Alignment and Red-teaming Engine for Multi-modal Toxicity Attack
von: Mao, Xutao, et al.
Veröffentlicht: (2026) -
Atoxia: Red-teaming Large Language Models with Target Toxic Answers
von: Du, Yuhao, et al.
Veröffentlicht: (2024)