Self-Disguise Attack: Induce the LLM to disguise itself for AIGT detection evasion
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Yinghan, Wen, Juan, Peng, Wanli, Wu, Zhengxian, Zhang, Ziwei, Xue, Yiming |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BadApex: Backdoor Attack Based on Adaptive Optimization Mechanism of Black-box Large Language Models
by: Wu, Zhengxian, et al.
Published: (2025)
by: Wu, Zhengxian, et al.
Published: (2025)
Is Your Writing Being Mimicked by AI? Unveiling Imitation with Invisible Watermarks in Creative Writing
by: Zhang, Ziwei, et al.
Published: (2025)
by: Zhang, Ziwei, et al.
Published: (2025)
GTSD: Generative Text Steganography Based on Diffusion Model
by: Wu, Zhengxian, et al.
Published: (2025)
by: Wu, Zhengxian, et al.
Published: (2025)
BeDKD: Backdoor Defense Based on Directional Mapping Module and Adversarial Knowledge Distillation
by: Wu, Zhengxian, et al.
Published: (2025)
by: Wu, Zhengxian, et al.
Published: (2025)
SLIP: Soft Label Mechanism and Key-Extraction-Guided CoT-based Defense Against Instruction Backdoor in APIs
by: Wu, Zhengxian, et al.
Published: (2025)
by: Wu, Zhengxian, et al.
Published: (2025)
Inhibitory Attacks on Backdoor-based Fingerprinting for Large Language Models
by: Fu, Hang, et al.
Published: (2026)
by: Fu, Hang, et al.
Published: (2026)
EditMF: Drawing an Invisible Fingerprint for Your Large Language Models
by: Wu, Jiaxuan, et al.
Published: (2025)
by: Wu, Jiaxuan, et al.
Published: (2025)
Kill two birds with one stone: generalized and robust AI-generated text detection via dynamic perturbations
by: Zhou, Yinghan, et al.
Published: (2025)
by: Zhou, Yinghan, et al.
Published: (2025)
Security Attacks on LLM-based Code Completion Tools
by: Cheng, Wen, et al.
Published: (2024)
by: Cheng, Wen, et al.
Published: (2024)
Are We in the AI-Generated Text World Already? Quantifying and Monitoring AIGT on Social Media
by: Sun, Zhen, et al.
Published: (2024)
by: Sun, Zhen, et al.
Published: (2024)
Retrieval-Confused Generation is a Good Defender for Privacy Violation Attack of Large Language Models
by: Peng, Wanli, et al.
Published: (2025)
by: Peng, Wanli, et al.
Published: (2025)
From Implicit to Explicit: Enhancing Self-Recognition in Large Language Models
by: Zhou, Yinghan, et al.
Published: (2025)
by: Zhou, Yinghan, et al.
Published: (2025)
Membership Inference Attacks Against In-Context Learning
by: Wen, Rui, et al.
Published: (2024)
by: Wen, Rui, et al.
Published: (2024)
LLM-Virus: Evolutionary Jailbreak Attack on Large Language Models
by: Yu, Miao, et al.
Published: (2024)
by: Yu, Miao, et al.
Published: (2024)
Defending LLM Watermarking Against Spoofing Attacks with Contrastive Representation Learning
by: An, Li, et al.
Published: (2025)
by: An, Li, et al.
Published: (2025)
GradingAttack: Exposing Security Vulnerabilities in LLM Based Educational Grading Agents
by: Li, Xueyi, et al.
Published: (2026)
by: Li, Xueyi, et al.
Published: (2026)
AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
by: Zeng, Yifan, et al.
Published: (2024)
by: Zeng, Yifan, et al.
Published: (2024)
Do Prompts Guarantee Safety? Mitigating Toxicity from LLM Generations through Subspace Intervention
by: Singh, Himanshu, et al.
Published: (2026)
by: Singh, Himanshu, et al.
Published: (2026)
Model-Agnostic Lifelong LLM Safety via Externalized Attack-Defense Co-Evolution
by: Zhang, Xiaozhe, et al.
Published: (2026)
by: Zhang, Xiaozhe, et al.
Published: (2026)
Semantic-Preserving Adversarial Attacks on LLMs: An Adaptive Greedy Binary Search Approach
by: Zhang, Chong, et al.
Published: (2025)
by: Zhang, Chong, et al.
Published: (2025)
PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks
by: Shen, Guobin, et al.
Published: (2025)
by: Shen, Guobin, et al.
Published: (2025)
Counterfactual Evaluation for Blind Attack Detection in LLM-based Evaluation Systems
by: Liu, Lijia, et al.
Published: (2025)
by: Liu, Lijia, et al.
Published: (2025)
Talk is (Not) Cheap: A Taxonomy and Benchmark Coverage Audit for LLM Attacks
by: Iyer, Karthik Raghu, et al.
Published: (2026)
by: Iyer, Karthik Raghu, et al.
Published: (2026)
FlexLLM: Exploring LLM Customization for Moving Target Defense on Black-Box LLMs Against Jailbreak Attacks
by: Chen, Bocheng, et al.
Published: (2024)
by: Chen, Bocheng, et al.
Published: (2024)
Proactive defense against LLM Jailbreak
by: Zhao, Weiliang, et al.
Published: (2025)
by: Zhao, Weiliang, et al.
Published: (2025)
Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreaking
by: Wu, Yu-Hang, et al.
Published: (2025)
by: Wu, Yu-Hang, et al.
Published: (2025)
An Anti-disguise Authentication System Using the First Impression of Avatar in Metaverse
by: Zhang, Zhenyong, et al.
Published: (2024)
by: Zhang, Zhenyong, et al.
Published: (2024)
Adversarial Attacks on LLM-as-a-Judge Systems: Insights from Prompt Injections
by: Maloyan, Narek, et al.
Published: (2025)
by: Maloyan, Narek, et al.
Published: (2025)
RedTWIZ: Diverse LLM Red Teaming via Adaptive Attack Planning
by: Horal, Artur, et al.
Published: (2025)
by: Horal, Artur, et al.
Published: (2025)
Network evasion detection with Bi-LSTM model
by: Chen, Kehua, et al.
Published: (2025)
by: Chen, Kehua, et al.
Published: (2025)
BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents
by: Wang, Yifei, et al.
Published: (2024)
by: Wang, Yifei, et al.
Published: (2024)
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense
by: Ouyang, Yang, et al.
Published: (2025)
by: Ouyang, Yang, et al.
Published: (2025)
Personalized Attacks of Social Engineering in Multi-turn Conversations: LLM Agents for Simulation and Detection
by: Kumarage, Tharindu, et al.
Published: (2025)
by: Kumarage, Tharindu, et al.
Published: (2025)
Data Extraction Attacks in Retrieval-Augmented Generation via Backdoors
by: Peng, Yuefeng, et al.
Published: (2024)
by: Peng, Yuefeng, et al.
Published: (2024)
Towards Label-Only Membership Inference Attack against Pre-trained Large Language Models
by: He, Yu, et al.
Published: (2025)
by: He, Yu, et al.
Published: (2025)
SecurityLingua: Efficient Defense of LLM Jailbreak Attacks via Security-Aware Prompt Compression
by: Li, Yucheng, et al.
Published: (2025)
by: Li, Yucheng, et al.
Published: (2025)
Transformers -- Messages in Disguise
by: Tyler, Joshua H., et al.
Published: (2024)
by: Tyler, Joshua H., et al.
Published: (2024)
Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges
by: Ding, Ruomeng, et al.
Published: (2026)
by: Ding, Ruomeng, et al.
Published: (2026)
"Give a Positive Review Only": An Early Investigation Into In-Paper Prompt Injection Attacks and Defenses for AI Reviewers
by: Zhou, Qin, et al.
Published: (2025)
by: Zhou, Qin, et al.
Published: (2025)
Crabs: Consuming Resource via Auto-generation for LLM-DoS Attack under Black-box Settings
by: Zhang, Yuanhe, et al.
Published: (2024)
by: Zhang, Yuanhe, et al.
Published: (2024)
Similar Items
-
BadApex: Backdoor Attack Based on Adaptive Optimization Mechanism of Black-box Large Language Models
by: Wu, Zhengxian, et al.
Published: (2025) -
Is Your Writing Being Mimicked by AI? Unveiling Imitation with Invisible Watermarks in Creative Writing
by: Zhang, Ziwei, et al.
Published: (2025) -
GTSD: Generative Text Steganography Based on Diffusion Model
by: Wu, Zhengxian, et al.
Published: (2025) -
BeDKD: Backdoor Defense Based on Directional Mapping Module and Adversarial Knowledge Distillation
by: Wu, Zhengxian, et al.
Published: (2025) -
SLIP: Soft Label Mechanism and Key-Extraction-Guided CoT-based Defense Against Instruction Backdoor in APIs
by: Wu, Zhengxian, et al.
Published: (2025)