Masked Language Model Based Textual Adversarial Example Detection
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zhang, Xiaomei, Zhang, Zhaoxi, Zhong, Qi, Zheng, Xufei, Zhang, Yanjun, Hu, Shengshan, Zhang, Leo Yu |
|---|---|
| Format: | Preprint |
| Publié: |
2023
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Large Language Model Watermark Stealing With Mixed Integer Programming
par: Zhang, Zhaoxi, et autres
Publié: (2024)
par: Zhang, Zhaoxi, et autres
Publié: (2024)
Exploring Gradient-Guided Masked Language Model to Detect Textual Adversarial Attacks
par: Zhang, Xiaomei, et autres
Publié: (2025)
par: Zhang, Xiaomei, et autres
Publié: (2025)
Less Is More -- Until It Breaks: Security Pitfalls of Vision Token Compression in Large Vision-Language Models
par: Zhang, Xiaomei, et autres
Publié: (2026)
par: Zhang, Xiaomei, et autres
Publié: (2026)
Character-Level Perturbations Disrupt LLM Watermarks
par: Zhang, Zhaoxi, et autres
Publié: (2025)
par: Zhang, Zhaoxi, et autres
Publié: (2025)
When Better Features Mean Greater Risks: The Performance-Privacy Trade-Off in Contrastive Learning
par: Sun, Ruining, et autres
Publié: (2025)
par: Sun, Ruining, et autres
Publié: (2025)
Towards Model Extraction Attacks in GAN-Based Image Translation via Domain Shift Mitigation
par: Mi, Di, et autres
Publié: (2024)
par: Mi, Di, et autres
Publié: (2024)
TEAM: Temporal Adversarial Examples Attack Model against Network Intrusion Detection System Applied to RNN
par: Liu, Ziyi, et autres
Publié: (2024)
par: Liu, Ziyi, et autres
Publié: (2024)
Kill Two Birds with One Stone! Trajectory enabled Unified Online Detection of Adversarial Examples and Backdoor Attacks
par: Fu, Anmin, et autres
Publié: (2025)
par: Fu, Anmin, et autres
Publié: (2025)
MARS: A Malignity-Aware Backdoor Defense in Federated Learning
par: Wan, Wei, et autres
Publié: (2025)
par: Wan, Wei, et autres
Publié: (2025)
From Pixels to Trajectory: Universal Adversarial Example Detection via Temporal Imprints
par: Gao, Yansong, et autres
Publié: (2025)
par: Gao, Yansong, et autres
Publié: (2025)
TaeBench: Improving Quality of Toxic Adversarial Examples
par: Zhu, Xuan, et autres
Publié: (2024)
par: Zhu, Xuan, et autres
Publié: (2024)
Mask-GCG: Are All Tokens in Adversarial Suffixes Necessary for Jailbreak Attacks?
par: Mu, Junjie, et autres
Publié: (2025)
par: Mu, Junjie, et autres
Publié: (2025)
A Real-Time, Self-Tuning Moderator Framework for Adversarial Prompt Detection
par: Zhang, Ivan
Publié: (2025)
par: Zhang, Ivan
Publié: (2025)
Creating Valid Adversarial Examples of Malware
par: Kozák, Matouš, et autres
Publié: (2023)
par: Kozák, Matouš, et autres
Publié: (2023)
Lightweight and Fast Backdoor Model Detection
par: Yu, Yinbo, et autres
Publié: (2026)
par: Yu, Yinbo, et autres
Publié: (2026)
ExplainableGuard: Interpretable Adversarial Defense for Large Language Models Using Chain-of-Thought Reasoning
par: Guan, Shaowei, et autres
Publié: (2025)
par: Guan, Shaowei, et autres
Publié: (2025)
Unseen Attack Detection in Software-Defined Networking Using a BERT-Based Large Language Model
par: Swileh, Mohammed N., et autres
Publié: (2024)
par: Swileh, Mohammed N., et autres
Publié: (2024)
Credential Leakage in LLM Agent Skills: A Large-Scale Empirical Study
par: Chen, Zhihao, et autres
Publié: (2026)
par: Chen, Zhihao, et autres
Publié: (2026)
Adversarial Attacks against Windows PE Malware Detection: A Survey of the State-of-the-Art
par: Ling, Xiang, et autres
Publié: (2021)
par: Ling, Xiang, et autres
Publié: (2021)
A Novel and Practical Universal Adversarial Perturbations against Deep Reinforcement Learning based Intrusion Detection Systems
par: Zhang, H., et autres
Publié: (2025)
par: Zhang, H., et autres
Publié: (2025)
Behind the Mask: Benchmarking Camouflaged Jailbreaks in Large Language Models
par: Zheng, Youjia, et autres
Publié: (2025)
par: Zheng, Youjia, et autres
Publié: (2025)
Improving Sustainability of Adversarial Examples in Class-Incremental Learning
par: Liu, Taifeng, et autres
Publié: (2025)
par: Liu, Taifeng, et autres
Publié: (2025)
NCCR: to Evaluate the Robustness of Neural Networks and Adversarial Examples
par: Pu, Shi, et autres
Publié: (2025)
par: Pu, Shi, et autres
Publié: (2025)
SNARE: Adaptive Scenario Synthesis for Eliciting Overeager Behavior in Coding Agents
par: Qu, Yubin, et autres
Publié: (2026)
par: Qu, Yubin, et autres
Publié: (2026)
Adversarial Evasion in Non-Stationary Malware Detection: Minimizing Drift Signals through Similarity-Constrained Perturbations
par: Acharya, Pawan, et autres
Publié: (2026)
par: Acharya, Pawan, et autres
Publié: (2026)
An Engorgio Prompt Makes Large Language Model Babble on
par: Dong, Jianshuo, et autres
Publié: (2024)
par: Dong, Jianshuo, et autres
Publié: (2024)
CL-Attack: Textual Backdoor Attacks via Cross-Lingual Triggers
par: Zheng, Jingyi, et autres
Publié: (2024)
par: Zheng, Jingyi, et autres
Publié: (2024)
Explainer-guided Targeted Adversarial Attacks against Binary Code Similarity Detection Models
par: Chen, Mingjie, et autres
Publié: (2025)
par: Chen, Mingjie, et autres
Publié: (2025)
State-Dependent Safety Failures in Multi-Turn Language Model Interaction
par: Li, Pengcheng, et autres
Publié: (2026)
par: Li, Pengcheng, et autres
Publié: (2026)
MEraser: An Effective Fingerprint Erasure Approach for Large Language Models
par: Zhang, Jingxuan, et autres
Publié: (2025)
par: Zhang, Jingxuan, et autres
Publié: (2025)
"Do Not Mention This to the User": Detecting and Understanding Malicious Agent Skills
par: Liu, Yi, et autres
Publié: (2026)
par: Liu, Yi, et autres
Publié: (2026)
Data-Free Model-Related Attacks: Unleashing the Potential of Generative AI
par: Ye, Dayong, et autres
Publié: (2025)
par: Ye, Dayong, et autres
Publié: (2025)
CodeBC: A More Secure Large Language Model for Smart Contract Code Generation in Blockchain
par: Wang, Lingxiang, et autres
Publié: (2025)
par: Wang, Lingxiang, et autres
Publié: (2025)
Membership Inference Attacks Against Vision-Language Models
par: Hu, Yuke, et autres
Publié: (2025)
par: Hu, Yuke, et autres
Publié: (2025)
Safeguarding Large Language Models: A Survey
par: Dong, Yi, et autres
Publié: (2024)
par: Dong, Yi, et autres
Publié: (2024)
Prefix Probing: Lightweight Harmful Content Detection for Large Language Models
par: Yang, Jirui, et autres
Publié: (2025)
par: Yang, Jirui, et autres
Publié: (2025)
Attention Masks Help Adversarial Attacks to Bypass Safety Detectors
par: Shi, Yunfan
Publié: (2024)
par: Shi, Yunfan
Publié: (2024)
SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models
par: Chen, Yulong, et autres
Publié: (2025)
par: Chen, Yulong, et autres
Publié: (2025)
VMask: Tunable Label Privacy Protection for Vertical Federated Learning via Layer Masking
par: Tan, Juntao, et autres
Publié: (2025)
par: Tan, Juntao, et autres
Publié: (2025)
MASKDROID: Robust Android Malware Detection with Masked Graph Representations
par: Zheng, Jingnan, et autres
Publié: (2024)
par: Zheng, Jingnan, et autres
Publié: (2024)
Documents similaires
-
Large Language Model Watermark Stealing With Mixed Integer Programming
par: Zhang, Zhaoxi, et autres
Publié: (2024) -
Exploring Gradient-Guided Masked Language Model to Detect Textual Adversarial Attacks
par: Zhang, Xiaomei, et autres
Publié: (2025) -
Less Is More -- Until It Breaks: Security Pitfalls of Vision Token Compression in Large Vision-Language Models
par: Zhang, Xiaomei, et autres
Publié: (2026) -
Character-Level Perturbations Disrupt LLM Watermarks
par: Zhang, Zhaoxi, et autres
Publié: (2025) -
When Better Features Mean Greater Risks: The Performance-Privacy Trade-Off in Contrastive Learning
par: Sun, Ruining, et autres
Publié: (2025)