STARE: Step-wise Temporal Alignment and Red-teaming Engine for Multi-modal Toxicity Attack
Fuente:
arXiv
Salvato in:
| Autori principali: | Mao, Xutao, Zhao, Liangjie, Liu, Tao, Zheng, Xiang, Zan, Hongying, Wang, Cong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Red-teaming the Multimodal Reasoning: Jailbreaking Vision-Language Models via Cross-modal Entanglement Attacks
di: Yan, Yu, et al.
Pubblicazione: (2026)
di: Yan, Yu, et al.
Pubblicazione: (2026)
DREAM: Dynamic Red-teaming across Environments for AI Models
di: Lu, Liming, et al.
Pubblicazione: (2025)
di: Lu, Liming, et al.
Pubblicazione: (2025)
Atoxia: Red-teaming Large Language Models with Target Toxic Answers
di: Du, Yuhao, et al.
Pubblicazione: (2024)
di: Du, Yuhao, et al.
Pubblicazione: (2024)
MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teaming
di: Guo, Weiyang, et al.
Pubblicazione: (2025)
di: Guo, Weiyang, et al.
Pubblicazione: (2025)
Temporal Logic-Based Multi-Vehicle Backdoor Attacks against Offline RL Agents in End-to-end Autonomous Driving
di: Chen, Xuan, et al.
Pubblicazione: (2025)
di: Chen, Xuan, et al.
Pubblicazione: (2025)
Red-Teaming LLM Multi-Agent Systems via Communication Attacks
di: He, Pengfei, et al.
Pubblicazione: (2025)
di: He, Pengfei, et al.
Pubblicazione: (2025)
Systematic Categorization, Construction and Evaluation of New Attacks against Multi-modal Mobile GUI Agents
di: Yang, Yulong, et al.
Pubblicazione: (2024)
di: Yang, Yulong, et al.
Pubblicazione: (2024)
An In-kernel Forensics Engine for Investigating Evasive Attacks
di: Zandi, Javad, et al.
Pubblicazione: (2025)
di: Zandi, Javad, et al.
Pubblicazione: (2025)
AlignSentinel: Alignment-Aware Detection of Prompt Injection Attacks
di: Jia, Yuqi, et al.
Pubblicazione: (2026)
di: Jia, Yuqi, et al.
Pubblicazione: (2026)
AJAR: Adaptive Jailbreak Architecture for Red-teaming
di: Dou, Yipu, et al.
Pubblicazione: (2026)
di: Dou, Yipu, et al.
Pubblicazione: (2026)
TRIDENT: Tri-modal Real-time Intrusion Detection Engine for New Targets
di: Alla, Ildi, et al.
Pubblicazione: (2025)
di: Alla, Ildi, et al.
Pubblicazione: (2025)
Introducing a New Alert Data Set for Multi-Step Attack Analysis
di: Landauer, Max, et al.
Pubblicazione: (2023)
di: Landauer, Max, et al.
Pubblicazione: (2023)
BadToken: Token-level Backdoor Attacks to Multi-modal Large Language Models
di: Yuan, Zenghui, et al.
Pubblicazione: (2025)
di: Yuan, Zenghui, et al.
Pubblicazione: (2025)
Impart: An Imperceptible and Effective Label-Specific Backdoor Attack
di: Zhao, Jingke, et al.
Pubblicazione: (2024)
di: Zhao, Jingke, et al.
Pubblicazione: (2024)
Arondight: Red Teaming Large Vision Language Models with Auto-generated Multi-modal Jailbreak Prompts
di: Liu, Yi, et al.
Pubblicazione: (2024)
di: Liu, Yi, et al.
Pubblicazione: (2024)
AdvAgent: Controllable Blackbox Red-teaming on Web Agents
di: Xu, Chejian, et al.
Pubblicazione: (2024)
di: Xu, Chejian, et al.
Pubblicazione: (2024)
SINCon: Mitigate LLM-Generated Malicious Message Injection Attack for Rumor Detection
di: Zhang, Mingqing, et al.
Pubblicazione: (2025)
di: Zhang, Mingqing, et al.
Pubblicazione: (2025)
ASTRA: Autonomous Spatial-Temporal Red-teaming for AI Software Assistants
di: Xu, Xiangzhe, et al.
Pubblicazione: (2025)
di: Xu, Xiangzhe, et al.
Pubblicazione: (2025)
Gotta Detect 'Em All: Fake Base Station and Multi-Step Attack Detection in Cellular Networks
di: Mubasshir, Kazi Samin, et al.
Pubblicazione: (2024)
di: Mubasshir, Kazi Samin, et al.
Pubblicazione: (2024)
SAGE: Sample-Aware Guarding Engine for Robust Intrusion Detection Against Adversarial Attacks
di: Chen, Jing, et al.
Pubblicazione: (2025)
di: Chen, Jing, et al.
Pubblicazione: (2025)
SkillAttack: Automated Red Teaming of Agent Skills through Attack Path Refinement
di: Duan, Zenghao, et al.
Pubblicazione: (2026)
di: Duan, Zenghao, et al.
Pubblicazione: (2026)
ProvAgent: Threat Detection Based on Identity-Behavior Binding and Multi-Agent Collaborative Attack Investigation
di: Yan, Wenhao, et al.
Pubblicazione: (2026)
di: Yan, Wenhao, et al.
Pubblicazione: (2026)
CL-Attack: Textual Backdoor Attacks via Cross-Lingual Triggers
di: Zheng, Jingyi, et al.
Pubblicazione: (2024)
di: Zheng, Jingyi, et al.
Pubblicazione: (2024)
AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases
di: Chen, Zhaorun, et al.
Pubblicazione: (2024)
di: Chen, Zhaorun, et al.
Pubblicazione: (2024)
Practical Reasoning Interruption Attacks on Reasoning Large Language Models
di: Cui, Yu, et al.
Pubblicazione: (2025)
di: Cui, Yu, et al.
Pubblicazione: (2025)
Typographic Attacks in a Multi-Image Setting
di: Wang, Xiaomeng, et al.
Pubblicazione: (2025)
di: Wang, Xiaomeng, et al.
Pubblicazione: (2025)
A Reward-driven Automated Webshell Malicious-code Generator for Red-teaming
di: Ding, Yizhong
Pubblicazione: (2025)
di: Ding, Yizhong
Pubblicazione: (2025)
ART: Automatic Red-teaming for Text-to-Image Models to Protect Benign Users
di: Li, Guanlin, et al.
Pubblicazione: (2024)
di: Li, Guanlin, et al.
Pubblicazione: (2024)
AgentVigil: Generic Black-Box Red-teaming for Indirect Prompt Injection against LLM Agents
di: Wang, Zhun, et al.
Pubblicazione: (2025)
di: Wang, Zhun, et al.
Pubblicazione: (2025)
Hijacking Attacks against Neural Networks by Analyzing Training Data
di: Ge, Yunjie, et al.
Pubblicazione: (2024)
di: Ge, Yunjie, et al.
Pubblicazione: (2024)
Phishing Detection in Ethereum via Temporal Graph Contrastive Learning
di: Wu, Cong, et al.
Pubblicazione: (2026)
di: Wu, Cong, et al.
Pubblicazione: (2026)
Benchmarking Fake Voice Detection in the Fake Voice Generation Arms Race
di: Mao, Xutao, et al.
Pubblicazione: (2025)
di: Mao, Xutao, et al.
Pubblicazione: (2025)
Cuckoo Attack: Stealthy and Persistent Attacks Against AI-IDE
di: Liu, Xinpeng, et al.
Pubblicazione: (2025)
di: Liu, Xinpeng, et al.
Pubblicazione: (2025)
Delayed Backdoor Attacks: Exploring the Temporal Dimension as a New Attack Surface in Pre-Trained Models
di: Ding, Zikang, et al.
Pubblicazione: (2026)
di: Ding, Zikang, et al.
Pubblicazione: (2026)
Non-Linear Trajectory Modeling for Multi-Step Gradient Inversion Attacks in Federated Learning
di: Xia, Li, et al.
Pubblicazione: (2025)
di: Xia, Li, et al.
Pubblicazione: (2025)
LaSM: Layer-wise Scaling Mechanism for Defending Pop-up Attack on GUI Agents
di: Yan, Zihe, et al.
Pubblicazione: (2025)
di: Yan, Zihe, et al.
Pubblicazione: (2025)
WFCAT: Augmenting Website Fingerprinting with Channel-wise Attention on Timing Features
di: Gong, Jiajun, et al.
Pubblicazione: (2024)
di: Gong, Jiajun, et al.
Pubblicazione: (2024)
SafeAligner: Safety Alignment against Jailbreak Attacks via Response Disparity Guidance
di: Huang, Caishuang, et al.
Pubblicazione: (2024)
di: Huang, Caishuang, et al.
Pubblicazione: (2024)
Rethinking Membership Inference Attacks Against Transfer Learning
di: Wu, Cong, et al.
Pubblicazione: (2025)
di: Wu, Cong, et al.
Pubblicazione: (2025)
LeakAgent: RL-based Red-teaming Agent for LLM Privacy Leakage
di: Nie, Yuzhou, et al.
Pubblicazione: (2024)
di: Nie, Yuzhou, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Red-teaming the Multimodal Reasoning: Jailbreaking Vision-Language Models via Cross-modal Entanglement Attacks
di: Yan, Yu, et al.
Pubblicazione: (2026) -
DREAM: Dynamic Red-teaming across Environments for AI Models
di: Lu, Liming, et al.
Pubblicazione: (2025) -
Atoxia: Red-teaming Large Language Models with Target Toxic Answers
di: Du, Yuhao, et al.
Pubblicazione: (2024) -
MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teaming
di: Guo, Weiyang, et al.
Pubblicazione: (2025) -
Temporal Logic-Based Multi-Vehicle Backdoor Attacks against Offline RL Agents in End-to-end Autonomous Driving
di: Chen, Xuan, et al.
Pubblicazione: (2025)