Exploring Gradient-Guided Masked Language Model to Detect Textual Adversarial Attacks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Xiaomei, Zhang, Zhaoxi, Zhang, Yanjun, Zheng, Xufei, Zhang, Leo Yu, Hu, Shengshan, Pan, Shirui |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Masked Language Model Based Textual Adversarial Example Detection
von: Zhang, Xiaomei, et al.
Veröffentlicht: (2023)
von: Zhang, Xiaomei, et al.
Veröffentlicht: (2023)
Large Language Model Watermark Stealing With Mixed Integer Programming
von: Zhang, Zhaoxi, et al.
Veröffentlicht: (2024)
von: Zhang, Zhaoxi, et al.
Veröffentlicht: (2024)
Less Is More -- Until It Breaks: Security Pitfalls of Vision Token Compression in Large Vision-Language Models
von: Zhang, Xiaomei, et al.
Veröffentlicht: (2026)
von: Zhang, Xiaomei, et al.
Veröffentlicht: (2026)
BiMark: Unbiased Multilayer Watermarking for Large Language Models
von: Feng, Xiaoyan, et al.
Veröffentlicht: (2025)
von: Feng, Xiaoyan, et al.
Veröffentlicht: (2025)
Character-Level Perturbations Disrupt LLM Watermarks
von: Zhang, Zhaoxi, et al.
Veröffentlicht: (2025)
von: Zhang, Zhaoxi, et al.
Veröffentlicht: (2025)
SEP-Attack: A Simple and Effective Paradigm for Transfer-Based Textual Adversarial Attack
von: Liu, Han, et al.
Veröffentlicht: (2026)
von: Liu, Han, et al.
Veröffentlicht: (2026)
Mask-GCG: Are All Tokens in Adversarial Suffixes Necessary for Jailbreak Attacks?
von: Mu, Junjie, et al.
Veröffentlicht: (2025)
von: Mu, Junjie, et al.
Veröffentlicht: (2025)
Fast Adversarial Training against Textual Adversarial Attacks
von: Yang, Yichen, et al.
Veröffentlicht: (2024)
von: Yang, Yichen, et al.
Veröffentlicht: (2024)
Breaking the Reviewer: Assessing the Vulnerability of Large Language Models in Automated Peer Review Under Textual Adversarial Attacks
von: Lin, Tzu-Ling, et al.
Veröffentlicht: (2025)
von: Lin, Tzu-Ling, et al.
Veröffentlicht: (2025)
Gradient Cuff: Detecting Jailbreak Attacks on Large Language Models by Exploring Refusal Loss Landscapes
von: Hu, Xiaomeng, et al.
Veröffentlicht: (2024)
von: Hu, Xiaomeng, et al.
Veröffentlicht: (2024)
HQA-Attack: Toward High Quality Black-Box Hard-Label Adversarial Attack on Text
von: Liu, Han, et al.
Veröffentlicht: (2024)
von: Liu, Han, et al.
Veröffentlicht: (2024)
LLM-Virus: Evolutionary Jailbreak Attack on Large Language Models
von: Yu, Miao, et al.
Veröffentlicht: (2024)
von: Yu, Miao, et al.
Veröffentlicht: (2024)
ELAD: Explanation-Guided Large Language Models Active Distillation
von: Zhang, Yifei, et al.
Veröffentlicht: (2024)
von: Zhang, Yifei, et al.
Veröffentlicht: (2024)
Knowledge Reasoning Language Model: Unifying Knowledge and Language for Inductive Knowledge Graph Reasoning
von: Zhuo, Xingrui, et al.
Veröffentlicht: (2025)
von: Zhuo, Xingrui, et al.
Veröffentlicht: (2025)
PB-UAP: Hybrid Universal Adversarial Attack For Image Segmentation
von: Song, Yufei, et al.
Veröffentlicht: (2024)
von: Song, Yufei, et al.
Veröffentlicht: (2024)
Emoti-Attack: Zero-Perturbation Adversarial Attacks on NLP Systems via Emoji Sequences
von: Zhang, Yangshijie
Veröffentlicht: (2025)
von: Zhang, Yangshijie
Veröffentlicht: (2025)
Claim-Guided Textual Backdoor Attack for Practical Applications
von: Song, Minkyoo, et al.
Veröffentlicht: (2024)
von: Song, Minkyoo, et al.
Veröffentlicht: (2024)
Less is More: Understanding Word-level Textual Adversarial Attack via n-gram Frequency Descend
von: Lu, Ning, et al.
Veröffentlicht: (2023)
von: Lu, Ning, et al.
Veröffentlicht: (2023)
SynDec: A Synthesize-then-Decode Approach for Arbitrary Textual Style Transfer via Large Language Models
von: Sun, Han, et al.
Veröffentlicht: (2025)
von: Sun, Han, et al.
Veröffentlicht: (2025)
TEG-DB: A Comprehensive Dataset and Benchmark of Textual-Edge Graphs
von: Li, Zhuofeng, et al.
Veröffentlicht: (2024)
von: Li, Zhuofeng, et al.
Veröffentlicht: (2024)
Fine-tuning Language Models with Generative Adversarial Reward Modelling
von: Yu, Zhang Ze, et al.
Veröffentlicht: (2023)
von: Yu, Zhang Ze, et al.
Veröffentlicht: (2023)
SPG: Sandwiched Policy Gradient for Masked Diffusion Language Models
von: Wang, Chenyu, et al.
Veröffentlicht: (2025)
von: Wang, Chenyu, et al.
Veröffentlicht: (2025)
Reasoning on Graphs: Faithful and Interpretable Large Language Model Reasoning
von: Luo, Linhao, et al.
Veröffentlicht: (2023)
von: Luo, Linhao, et al.
Veröffentlicht: (2023)
A Trembling House of Cards? Mapping Adversarial Attacks against Language Agents
von: Mo, Lingbo, et al.
Veröffentlicht: (2024)
von: Mo, Lingbo, et al.
Veröffentlicht: (2024)
When Better Features Mean Greater Risks: The Performance-Privacy Trade-Off in Contrastive Learning
von: Sun, Ruining, et al.
Veröffentlicht: (2025)
von: Sun, Ruining, et al.
Veröffentlicht: (2025)
SuperMerge: An Approach For Gradient-Based Model Merging
von: Yang, Haoyu, et al.
Veröffentlicht: (2024)
von: Yang, Haoyu, et al.
Veröffentlicht: (2024)
UniGuardian: A Unified Defense for Detecting Prompt Injection, Backdoor Attacks and Adversarial Attacks in Large Language Models
von: Lin, Huawei, et al.
Veröffentlicht: (2025)
von: Lin, Huawei, et al.
Veröffentlicht: (2025)
Abstraction-of-Thought Makes Language Models Better Reasoners
von: Hong, Ruixin, et al.
Veröffentlicht: (2024)
von: Hong, Ruixin, et al.
Veröffentlicht: (2024)
Dial-MAE: ConTextual Masked Auto-Encoder for Retrieval-based Dialogue Systems
von: Su, Zhenpeng, et al.
Veröffentlicht: (2023)
von: Su, Zhenpeng, et al.
Veröffentlicht: (2023)
"Do Not Mention This to the User": Detecting and Understanding Malicious Agent Skills
von: Liu, Yi, et al.
Veröffentlicht: (2026)
von: Liu, Yi, et al.
Veröffentlicht: (2026)
Masked Diffusion Models are Secretly Time-Agnostic Masked Models and Exploit Inaccurate Categorical Sampling
von: Zheng, Kaiwen, et al.
Veröffentlicht: (2024)
von: Zheng, Kaiwen, et al.
Veröffentlicht: (2024)
Enhancing Large Language Model for Knowledge Graph Completion via Structure-Aware Alignment-Tuning
von: Liu, Yu, et al.
Veröffentlicht: (2025)
von: Liu, Yu, et al.
Veröffentlicht: (2025)
Boosting Large Language Models with Mask Fine-Tuning
von: Zhang, Mingyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Mingyuan, et al.
Veröffentlicht: (2025)
MultiAgent Collaboration Attack: Investigating Adversarial Attacks in Large Language Model Collaborations via Debate
von: Amayuelas, Alfonso, et al.
Veröffentlicht: (2024)
von: Amayuelas, Alfonso, et al.
Veröffentlicht: (2024)
SNARE: Adaptive Scenario Synthesis for Eliciting Overeager Behavior in Coding Agents
von: Qu, Yubin, et al.
Veröffentlicht: (2026)
von: Qu, Yubin, et al.
Veröffentlicht: (2026)
Exploring and Evaluating Multimodal Knowledge Reasoning Consistency of Multimodal Large Language Models
von: Jia, Boyu, et al.
Veröffentlicht: (2025)
von: Jia, Boyu, et al.
Veröffentlicht: (2025)
Integrating Large Language Models and Knowledge Graphs for Extraction and Validation of Textual Test Data
von: De Santis, Antonio, et al.
Veröffentlicht: (2024)
von: De Santis, Antonio, et al.
Veröffentlicht: (2024)
Exploring the Compositional Deficiency of Large Language Models in Mathematical Reasoning
von: Zhao, Jun, et al.
Veröffentlicht: (2024)
von: Zhao, Jun, et al.
Veröffentlicht: (2024)
Exploring Polyglot Harmony: On Multilingual Data Allocation for Large Language Models Pretraining
von: Guo, Ping, et al.
Veröffentlicht: (2025)
von: Guo, Ping, et al.
Veröffentlicht: (2025)
$\text{M}^{2}$LLM: Multi-view Molecular Representation Learning with Large Language Models
von: Ju, Jiaxin, et al.
Veröffentlicht: (2025)
von: Ju, Jiaxin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Masked Language Model Based Textual Adversarial Example Detection
von: Zhang, Xiaomei, et al.
Veröffentlicht: (2023) -
Large Language Model Watermark Stealing With Mixed Integer Programming
von: Zhang, Zhaoxi, et al.
Veröffentlicht: (2024) -
Less Is More -- Until It Breaks: Security Pitfalls of Vision Token Compression in Large Vision-Language Models
von: Zhang, Xiaomei, et al.
Veröffentlicht: (2026) -
BiMark: Unbiased Multilayer Watermarking for Large Language Models
von: Feng, Xiaoyan, et al.
Veröffentlicht: (2025) -
Character-Level Perturbations Disrupt LLM Watermarks
von: Zhang, Zhaoxi, et al.
Veröffentlicht: (2025)