DualSentinel: A Lightweight Framework for Detecting Targeted Attacks in Black-box LLM via Dual Entropy Lull Pattern
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pang, Xiaoyi, Hao, Xuanyi, Liu, Pengyu, Luo, Qi, Guo, Song, Wang, Zhibo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Safeguarding LLM Embeddings in End-Cloud Collaboration via Entropy-Driven Perturbation
von: Jin, Shuaifan, et al.
Veröffentlicht: (2025)
von: Jin, Shuaifan, et al.
Veröffentlicht: (2025)
Black-box Membership Inference Attacks against Fine-tuned Diffusion Models
von: Pang, Yan, et al.
Veröffentlicht: (2023)
von: Pang, Yan, et al.
Veröffentlicht: (2023)
DV-FSR: A Dual-View Target Attack Framework for Federated Sequential Recommendation
von: Qin, Qitao, et al.
Veröffentlicht: (2024)
von: Qin, Qitao, et al.
Veröffentlicht: (2024)
WADBERT: Dual-channel Web Attack Detection Based on BERT Models
von: Luo, Kangqiang, et al.
Veröffentlicht: (2026)
von: Luo, Kangqiang, et al.
Veröffentlicht: (2026)
AlignSentinel: Alignment-Aware Detection of Prompt Injection Attacks
von: Jia, Yuqi, et al.
Veröffentlicht: (2026)
von: Jia, Yuqi, et al.
Veröffentlicht: (2026)
Mind the Gap: Detecting Black-box Adversarial Attacks in the Making through Query Update Analysis
von: Park, Jeonghwan, et al.
Veröffentlicht: (2025)
von: Park, Jeonghwan, et al.
Veröffentlicht: (2025)
When Reasoning Leaks Membership: Membership Inference Attack on Black-box Large Reasoning Models
von: Hu, Ruihan, et al.
Veröffentlicht: (2026)
von: Hu, Ruihan, et al.
Veröffentlicht: (2026)
A Wolf in Sheep's Clothing: Practical Black-box Adversarial Attacks for Evading Learning-based Windows Malware Detection in the Wild
von: Ling, Xiang, et al.
Veröffentlicht: (2024)
von: Ling, Xiang, et al.
Veröffentlicht: (2024)
White-box Membership Inference Attacks against Diffusion Models
von: Pang, Yan, et al.
Veröffentlicht: (2023)
von: Pang, Yan, et al.
Veröffentlicht: (2023)
MalModel: Hiding Malicious Payload in Mobile Deep Learning Models with Black-box Backdoor Attack
von: Hua, Jiayi, et al.
Veröffentlicht: (2024)
von: Hua, Jiayi, et al.
Veröffentlicht: (2024)
Practicable Black-box Evasion Attacks on Link Prediction in Dynamic Graphs -- A Graph Sequential Embedding Method
von: Li, Jiate, et al.
Veröffentlicht: (2024)
von: Li, Jiate, et al.
Veröffentlicht: (2024)
BlackboxBench: A Comprehensive Benchmark of Black-box Adversarial Attacks
von: Zheng, Meixi, et al.
Veröffentlicht: (2023)
von: Zheng, Meixi, et al.
Veröffentlicht: (2023)
Pop Quiz Attack: Black-box Membership Inference Attacks Against Large Language Models
von: Chen, Zeyuan, et al.
Veröffentlicht: (2026)
von: Chen, Zeyuan, et al.
Veröffentlicht: (2026)
DataSentinel: A Game-Theoretic Detection of Prompt Injection Attacks
von: Liu, Yupei, et al.
Veröffentlicht: (2025)
von: Liu, Yupei, et al.
Veröffentlicht: (2025)
Select Me! When You Need a Tool: A Black-box Text Attack on Tool Selection
von: Chen, Liuji, et al.
Veröffentlicht: (2025)
von: Chen, Liuji, et al.
Veröffentlicht: (2025)
Discovering New Shadow Patterns for Black-Box Attacks on Lane Detection of Autonomous Vehicles
von: MohajerAnsari, Pedram, et al.
Veröffentlicht: (2024)
von: MohajerAnsari, Pedram, et al.
Veröffentlicht: (2024)
RL-JACK: Reinforcement Learning-powered Black-box Jailbreaking Attack against LLMs
von: Chen, Xuan, et al.
Veröffentlicht: (2024)
von: Chen, Xuan, et al.
Veröffentlicht: (2024)
Multi-task Adversarial Attacks against Black-box Model with Few-shot Queries
von: Wang, Wenqiang, et al.
Veröffentlicht: (2025)
von: Wang, Wenqiang, et al.
Veröffentlicht: (2025)
Lite-BD: A Lightweight Black-box Backdoor Defense via Reviving Multi-Stage Image Transformations
von: Miah, Abdullah Arafat, et al.
Veröffentlicht: (2026)
von: Miah, Abdullah Arafat, et al.
Veröffentlicht: (2026)
TPPR: APT Tactic / Technique Pattern Guided Attack Path Reasoning for Attack Investigation
von: Sheng, Qi
Veröffentlicht: (2025)
von: Sheng, Qi
Veröffentlicht: (2025)
FBA$^2$D: Frequency-based Black-box Attack for AI-generated Image Detection
von: Chen, Xiaojing, et al.
Veröffentlicht: (2025)
von: Chen, Xiaojing, et al.
Veröffentlicht: (2025)
MIRAGE: Misleading Retrieval-Augmented Generation via Black-box and Query-agnostic Poisoning Attacks
von: Chen, Tailun, et al.
Veröffentlicht: (2025)
von: Chen, Tailun, et al.
Veröffentlicht: (2025)
Query Provenance Analysis: Efficient and Robust Defense against Query-based Black-box Attacks
von: Li, Shaofei, et al.
Veröffentlicht: (2024)
von: Li, Shaofei, et al.
Veröffentlicht: (2024)
BETA: Automated Black-box Exploration for Timing Attacks in Processors
von: Chen, Congcong, et al.
Veröffentlicht: (2024)
von: Chen, Congcong, et al.
Veröffentlicht: (2024)
Dynamic Black-box Backdoor Attacks on IoT Sensory Data
von: Chathoth, Ajesh Koyatan, et al.
Veröffentlicht: (2025)
von: Chathoth, Ajesh Koyatan, et al.
Veröffentlicht: (2025)
Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection
von: Koide, Takashi, et al.
Veröffentlicht: (2026)
von: Koide, Takashi, et al.
Veröffentlicht: (2026)
VisInject: Disruption != Injection -- A Dual-Dimension Evaluation of Universal Adversarial Attacks on Vision-Language Models
von: Liu, Pang, et al.
Veröffentlicht: (2026)
von: Liu, Pang, et al.
Veröffentlicht: (2026)
WebSentinel: Detecting and Localizing Prompt Injection Attacks for Web Agents
von: Wang, Xilong, et al.
Veröffentlicht: (2026)
von: Wang, Xilong, et al.
Veröffentlicht: (2026)
KG-DF: A Black-box Defense Framework against Jailbreak Attacks Based on Knowledge Graphs
von: Liu, Shuyuan, et al.
Veröffentlicht: (2025)
von: Liu, Shuyuan, et al.
Veröffentlicht: (2025)
From Head to Tail: Efficient Black-box Model Inversion Attack via Long-tailed Learning
von: Li, Ziang, et al.
Veröffentlicht: (2025)
von: Li, Ziang, et al.
Veröffentlicht: (2025)
EvadeDroid: A Practical Evasion Attack on Machine Learning for Black-box Android Malware Detection
von: Bostani, Hamid, et al.
Veröffentlicht: (2021)
von: Bostani, Hamid, et al.
Veröffentlicht: (2021)
Collaborative Shadows: Distributed Backdoor Attacks in LLM-Based Multi-Agent Systems
von: Zhu, Pengyu, et al.
Veröffentlicht: (2025)
von: Zhu, Pengyu, et al.
Veröffentlicht: (2025)
SEA: Shareable and Explainable Attribution for Query-based Black-box Attacks
von: Gao, Yue, et al.
Veröffentlicht: (2023)
von: Gao, Yue, et al.
Veröffentlicht: (2023)
Black-box Optimization of LLM Outputs by Asking for Directions
von: Zhang, Jie, et al.
Veröffentlicht: (2025)
von: Zhang, Jie, et al.
Veröffentlicht: (2025)
Targeted Adversarial Traffic Generation : Black-box Approach to Evade Intrusion Detection Systems in IoT Networks
von: Debicha, Islam, et al.
Veröffentlicht: (2026)
von: Debicha, Islam, et al.
Veröffentlicht: (2026)
DualGuard: Dual-stream Large Language Model Watermarking Defense against Paraphrase and Spoofing Attack
von: Li, Hao, et al.
Veröffentlicht: (2025)
von: Li, Hao, et al.
Veröffentlicht: (2025)
JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks
von: Zhang, Xiaoyu, et al.
Veröffentlicht: (2023)
von: Zhang, Xiaoyu, et al.
Veröffentlicht: (2023)
Online Poisoning Attack Against Reinforcement Learning under Black-box Environments
von: Li, Jianhui, et al.
Veröffentlicht: (2024)
von: Li, Jianhui, et al.
Veröffentlicht: (2024)
DualBreach: Efficient Dual-Jailbreaking via Target-Driven Initialization and Multi-Target Optimization
von: Huang, Xinzhe, et al.
Veröffentlicht: (2025)
von: Huang, Xinzhe, et al.
Veröffentlicht: (2025)
Performance-lossless Black-box Model Watermarking
von: Zhao, Na, et al.
Veröffentlicht: (2023)
von: Zhao, Na, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Safeguarding LLM Embeddings in End-Cloud Collaboration via Entropy-Driven Perturbation
von: Jin, Shuaifan, et al.
Veröffentlicht: (2025) -
Black-box Membership Inference Attacks against Fine-tuned Diffusion Models
von: Pang, Yan, et al.
Veröffentlicht: (2023) -
DV-FSR: A Dual-View Target Attack Framework for Federated Sequential Recommendation
von: Qin, Qitao, et al.
Veröffentlicht: (2024) -
WADBERT: Dual-channel Web Attack Detection Based on BERT Models
von: Luo, Kangqiang, et al.
Veröffentlicht: (2026) -
AlignSentinel: Alignment-Aware Detection of Prompt Injection Attacks
von: Jia, Yuqi, et al.
Veröffentlicht: (2026)