Co-Evolutionary Multi-Modal Alignment via Structured Adversarial Evolution
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shi, Guoxin, Wang, Haoyu, Yang, Zaihui, Wang, Yuxing, Chang, Yongzhe |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Probing the Safety Response Boundary of Large Language Models via Unsafe Decoding Path Generation
von: Wang, Haoyu, et al.
Veröffentlicht: (2024)
von: Wang, Haoyu, et al.
Veröffentlicht: (2024)
Adversarial Attack-Defense Co-Evolution for LLM Safety Alignment via Tree-Group Dual-Aware Search and Optimization
von: Li, Xurui, et al.
Veröffentlicht: (2025)
von: Li, Xurui, et al.
Veröffentlicht: (2025)
MoCo-EA: Exploiting Adversarial Mode Connectivity for Efficient Evolutionary Attacks
von: Kim, Hyo Seo, et al.
Veröffentlicht: (2026)
von: Kim, Hyo Seo, et al.
Veröffentlicht: (2026)
ShellForge: Adversarial Co-Evolution of Webshell Generation and Multi-View Detection for Robust Webshell Defense
von: Ding, Yizhong
Veröffentlicht: (2026)
von: Ding, Yizhong
Veröffentlicht: (2026)
Medical Multimodal Model Stealing Attacks via Adversarial Domain Alignment
von: Shen, Yaling, et al.
Veröffentlicht: (2025)
von: Shen, Yaling, et al.
Veröffentlicht: (2025)
Adversarial Illusions in Multi-Modal Embeddings
von: Zhang, Tingwei, et al.
Veröffentlicht: (2023)
von: Zhang, Tingwei, et al.
Veröffentlicht: (2023)
Magmaw: Modality-Agnostic Adversarial Attacks on Machine Learning-Based Wireless Communication Systems
von: Chang, Jung-Woo, et al.
Veröffentlicht: (2023)
von: Chang, Jung-Woo, et al.
Veröffentlicht: (2023)
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training
von: Du, Pengfei
Veröffentlicht: (2025)
von: Du, Pengfei
Veröffentlicht: (2025)
Knowledge Poisoning Attacks on Medical Multi-Modal Retrieval-Augmented Generation
von: Yang, Peiru, et al.
Veröffentlicht: (2026)
von: Yang, Peiru, et al.
Veröffentlicht: (2026)
NCCR: to Evaluate the Robustness of Neural Networks and Adversarial Examples
von: Pu, Shi, et al.
Veröffentlicht: (2025)
von: Pu, Shi, et al.
Veröffentlicht: (2025)
On the Feasibility of Using MultiModal LLMs to Execute AR Social Engineering Attacks
von: Bi, Ting, et al.
Veröffentlicht: (2025)
von: Bi, Ting, et al.
Veröffentlicht: (2025)
From Allies to Adversaries: Manipulating LLM Tool-Calling through Adversarial Injection
von: Wang, Haowei, et al.
Veröffentlicht: (2024)
von: Wang, Haowei, et al.
Veröffentlicht: (2024)
Visual-RolePlay: Universal Jailbreak Attack on MultiModal Large Language Models via Role-playing Image Character
von: Ma, Siyuan, et al.
Veröffentlicht: (2024)
von: Ma, Siyuan, et al.
Veröffentlicht: (2024)
Scam Shield: Multi-Model Voting and Fine-Tuned LLMs Against Adversarial Attacks
von: Chang, Chen-Wei, et al.
Veröffentlicht: (2025)
von: Chang, Chen-Wei, et al.
Veröffentlicht: (2025)
Agent Safety Alignment via Reinforcement Learning
von: Sha, Zeyang, et al.
Veröffentlicht: (2025)
von: Sha, Zeyang, et al.
Veröffentlicht: (2025)
Poster: Long PHP webshell files detection based on sliding window attention
von: Wang, Zhiqiang, et al.
Veröffentlicht: (2025)
von: Wang, Zhiqiang, et al.
Veröffentlicht: (2025)
Security-aware Semantic-driven ISAC via Paired Adversarial Residual Networks
von: Liu, Yu, et al.
Veröffentlicht: (2025)
von: Liu, Yu, et al.
Veröffentlicht: (2025)
MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teaming
von: Guo, Weiyang, et al.
Veröffentlicht: (2025)
von: Guo, Weiyang, et al.
Veröffentlicht: (2025)
EVA: Editing for Versatile Alignment against Jailbreaks
von: Wang, Yi, et al.
Veröffentlicht: (2026)
von: Wang, Yi, et al.
Veröffentlicht: (2026)
Multi-task Adversarial Attacks against Black-box Model with Few-shot Queries
von: Wang, Wenqiang, et al.
Veröffentlicht: (2025)
von: Wang, Wenqiang, et al.
Veröffentlicht: (2025)
Attention Masks Help Adversarial Attacks to Bypass Safety Detectors
von: Shi, Yunfan
Veröffentlicht: (2024)
von: Shi, Yunfan
Veröffentlicht: (2024)
BinCtx: Multi-Modal Representation Learning for Robust Android App Behavior Detection
von: Liu, Zichen, et al.
Veröffentlicht: (2025)
von: Liu, Zichen, et al.
Veröffentlicht: (2025)
CyberEvolver: Structured Self-Evolution for Cybersecurity Agents On the Fly
von: Fan, Yihe, et al.
Veröffentlicht: (2026)
von: Fan, Yihe, et al.
Veröffentlicht: (2026)
On the (In)Security of LLM App Stores
von: Hou, Xinyi, et al.
Veröffentlicht: (2024)
von: Hou, Xinyi, et al.
Veröffentlicht: (2024)
Multi-Stream Perturbation Attack: Breaking Safety Alignment of Thinking LLMs Through Concurrent Task Interference
von: Yang, Fan
Veröffentlicht: (2026)
von: Yang, Fan
Veröffentlicht: (2026)
Model Context Protocol (MCP): Landscape, Security Threats, and Future Research Directions
von: Hou, Xinyi, et al.
Veröffentlicht: (2025)
von: Hou, Xinyi, et al.
Veröffentlicht: (2025)
PiCo: Jailbreaking Multimodal Large Language Models via Pictorial Code Contextualization
von: Liu, Aofan, et al.
Veröffentlicht: (2025)
von: Liu, Aofan, et al.
Veröffentlicht: (2025)
AdaDoS: Adaptive DoS Attack via Deep Adversarial Reinforcement Learning in SDN
von: Shao, Wei, et al.
Veröffentlicht: (2025)
von: Shao, Wei, et al.
Veröffentlicht: (2025)
Omni-Safety under Cross-Modality Conflict: Vulnerabilities, Dynamics Mechanisms and Efficient Alignment
von: Wang, Kun, et al.
Veröffentlicht: (2026)
von: Wang, Kun, et al.
Veröffentlicht: (2026)
Evolutionary Trigger Detection and Lightweight Model Repair Based Backdoor Defense
von: Zhou, Qi, et al.
Veröffentlicht: (2024)
von: Zhou, Qi, et al.
Veröffentlicht: (2024)
Enhancing Adversarial Resistance in LLMs with Recursion
von: Li, Bryan, et al.
Veröffentlicht: (2024)
von: Li, Bryan, et al.
Veröffentlicht: (2024)
Multi-turn Jailbreaking Attack in Multi-Modal Large Language Models
von: Das, Badhan Chandra, et al.
Veröffentlicht: (2026)
von: Das, Badhan Chandra, et al.
Veröffentlicht: (2026)
From Pixels to Trajectory: Universal Adversarial Example Detection via Temporal Imprints
von: Gao, Yansong, et al.
Veröffentlicht: (2025)
von: Gao, Yansong, et al.
Veröffentlicht: (2025)
Topological Signatures of Adversaries in Multimodal Alignments
von: Vu, Minh, et al.
Veröffentlicht: (2025)
von: Vu, Minh, et al.
Veröffentlicht: (2025)
Frequency-Domain Regularized Adversarial Alignment for Transferable Attacks against Closed-Source MLLMs
von: Yuan, Leitao, et al.
Veröffentlicht: (2026)
von: Yuan, Leitao, et al.
Veröffentlicht: (2026)
From Assistants to Adversaries: Exploring the Security Risks of Mobile LLM Agents
von: Wu, Liangxuan, et al.
Veröffentlicht: (2025)
von: Wu, Liangxuan, et al.
Veröffentlicht: (2025)
Reflect-Guard: Enhancing LLM Safeguards against Adversarial Prompts via Logical Self-Reflection
von: Lin, Lixing, et al.
Veröffentlicht: (2026)
von: Lin, Lixing, et al.
Veröffentlicht: (2026)
Bergeron: Combating Adversarial Attacks through a Conscience-Based Alignment Framework
von: Pisano, Matthew, et al.
Veröffentlicht: (2023)
von: Pisano, Matthew, et al.
Veröffentlicht: (2023)
Improving Generalization on Cybersecurity Tasks with Multi-Modal Contrastive Learning
von: Huang, Jianan, et al.
Veröffentlicht: (2026)
von: Huang, Jianan, et al.
Veröffentlicht: (2026)
NetDeTox: Adversarial and Efficient Evasion of Hardware-Security GNNs via RL-LLM Orchestration
von: Wang, Zeng, et al.
Veröffentlicht: (2025)
von: Wang, Zeng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Probing the Safety Response Boundary of Large Language Models via Unsafe Decoding Path Generation
von: Wang, Haoyu, et al.
Veröffentlicht: (2024) -
Adversarial Attack-Defense Co-Evolution for LLM Safety Alignment via Tree-Group Dual-Aware Search and Optimization
von: Li, Xurui, et al.
Veröffentlicht: (2025) -
MoCo-EA: Exploiting Adversarial Mode Connectivity for Efficient Evolutionary Attacks
von: Kim, Hyo Seo, et al.
Veröffentlicht: (2026) -
ShellForge: Adversarial Co-Evolution of Webshell Generation and Multi-View Detection for Robust Webshell Defense
von: Ding, Yizhong
Veröffentlicht: (2026) -
Medical Multimodal Model Stealing Attacks via Adversarial Domain Alignment
von: Shen, Yaling, et al.
Veröffentlicht: (2025)