Humanizing the Machine: Proxy Attacks to Mislead LLM Detectors
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Tianchun, Chen, Yuanzhou, Liu, Zichuan, Chen, Zhanwen, Chen, Haifeng, Zhang, Xiang, Cheng, Wei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Synthetic Data Can Mislead Evaluations: Membership Inference as Machine Text Detection
von: Naseh, Ali, et al.
Veröffentlicht: (2025)
von: Naseh, Ali, et al.
Veröffentlicht: (2025)
MIA-Tuner: Adapting Large Language Models as Pre-training Text Detector
von: Fu, Wenjie, et al.
Veröffentlicht: (2024)
von: Fu, Wenjie, et al.
Veröffentlicht: (2024)
Humanizing Machine-Generated Content: Evading AI-Text Detection through Adversarial Attack
von: Zhou, Ying, et al.
Veröffentlicht: (2024)
von: Zhou, Ying, et al.
Veröffentlicht: (2024)
AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
von: Zeng, Yifan, et al.
Veröffentlicht: (2024)
von: Zeng, Yifan, et al.
Veröffentlicht: (2024)
GCG Attack On A Diffusion LLM
von: Neyroud, Ruben, et al.
Veröffentlicht: (2025)
von: Neyroud, Ruben, et al.
Veröffentlicht: (2025)
PAL: Proxy-Guided Black-Box Attack on Large Language Models
von: Sitawarin, Chawin, et al.
Veröffentlicht: (2024)
von: Sitawarin, Chawin, et al.
Veröffentlicht: (2024)
Can Federated Learning Safeguard Private Data in LLM Training? Vulnerabilities, Attacks, and Defense Evaluation
von: Guo, Wenkai, et al.
Veröffentlicht: (2025)
von: Guo, Wenkai, et al.
Veröffentlicht: (2025)
Practical Membership Inference Attacks against Fine-tuned Large Language Models via Self-prompt Calibration
von: Fu, Wenjie, et al.
Veröffentlicht: (2023)
von: Fu, Wenjie, et al.
Veröffentlicht: (2023)
Less is More: Understanding Word-level Textual Adversarial Attack via n-gram Frequency Descend
von: Lu, Ning, et al.
Veröffentlicht: (2023)
von: Lu, Ning, et al.
Veröffentlicht: (2023)
Protecting Your LLMs with Information Bottleneck
von: Liu, Zichuan, et al.
Veröffentlicht: (2024)
von: Liu, Zichuan, et al.
Veröffentlicht: (2024)
Reversible Jump Attack to Textual Classifiers with Modification Reduction
von: Ni, Mingze, et al.
Veröffentlicht: (2024)
von: Ni, Mingze, et al.
Veröffentlicht: (2024)
Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM
von: Cao, Bochuan, et al.
Veröffentlicht: (2023)
von: Cao, Bochuan, et al.
Veröffentlicht: (2023)
PVMark: Enabling Public Verifiability for LLM Watermarking Schemes
von: Duan, Haohua, et al.
Veröffentlicht: (2025)
von: Duan, Haohua, et al.
Veröffentlicht: (2025)
Cross-Entropy Attacks to Language Models via Rare Event Simulation
von: Ni, Mingze, et al.
Veröffentlicht: (2025)
von: Ni, Mingze, et al.
Veröffentlicht: (2025)
Cascading and Proxy Membership Inference Attacks
von: Du, Yuntao, et al.
Veröffentlicht: (2025)
von: Du, Yuntao, et al.
Veröffentlicht: (2025)
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval
von: Chen, Taiye, et al.
Veröffentlicht: (2025)
von: Chen, Taiye, et al.
Veröffentlicht: (2025)
MetaMoE: Diversity-Aware Proxy Selection for Privacy-Preserving Mixture-of-Experts Unification
von: Jiang, Weisen, et al.
Veröffentlicht: (2026)
von: Jiang, Weisen, et al.
Veröffentlicht: (2026)
Measuring Real-World Prompt Injection Attacks in LLM-based Resume Screening
von: Zhang, Mohan, et al.
Veröffentlicht: (2026)
von: Zhang, Mohan, et al.
Veröffentlicht: (2026)
Chain of Attack: a Semantic-Driven Contextual Multi-Turn attacker for LLM
von: Yang, Xikang, et al.
Veröffentlicht: (2024)
von: Yang, Xikang, et al.
Veröffentlicht: (2024)
Membership Inference Attacks on LLM-based Recommender Systems
von: He, Jiajie, et al.
Veröffentlicht: (2025)
von: He, Jiajie, et al.
Veröffentlicht: (2025)
Token-Modification Adversarial Attacks for Natural Language Processing: A Survey
von: Roth, Tom, et al.
Veröffentlicht: (2021)
von: Roth, Tom, et al.
Veröffentlicht: (2021)
Evaluations of Machine Learning Privacy Defenses are Misleading
von: Aerni, Michael, et al.
Veröffentlicht: (2024)
von: Aerni, Michael, et al.
Veröffentlicht: (2024)
A Framework for Cost-Effective and Self-Adaptive LLM Shaking and Recovery Mechanism
von: Chen, Zhiyu, et al.
Veröffentlicht: (2024)
von: Chen, Zhiyu, et al.
Veröffentlicht: (2024)
BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents
von: Wang, Yifei, et al.
Veröffentlicht: (2024)
von: Wang, Yifei, et al.
Veröffentlicht: (2024)
Why LLM Safety Guardrails Collapse After Fine-tuning: A Similarity Analysis Between Alignment and Fine-tuning Datasets
von: Hsiung, Lei, et al.
Veröffentlicht: (2025)
von: Hsiung, Lei, et al.
Veröffentlicht: (2025)
Eyes-on-Me: Scalable RAG Poisoning through Transferable Attention-Steering Attractors
von: Chen, Yen-Shan, et al.
Veröffentlicht: (2025)
von: Chen, Yen-Shan, et al.
Veröffentlicht: (2025)
Prompt2Fingerprint: Plug-and-Play LLM Fingerprinting via Text-to-Weight Generation
von: Chen, Sixu, et al.
Veröffentlicht: (2026)
von: Chen, Sixu, et al.
Veröffentlicht: (2026)
MPAT: Building Robust Deep Neural Networks against Textual Adversarial Attacks
von: Zhang, Fangyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Fangyuan, et al.
Veröffentlicht: (2024)
Membership Inference Attacks and Privacy in Topic Modeling
von: Manzonelli, Nico, et al.
Veröffentlicht: (2024)
von: Manzonelli, Nico, et al.
Veröffentlicht: (2024)
Fault Sneaking Attack: a Stealthy Framework for Misleading Deep Neural Networks
von: Zhao, Pu, et al.
Veröffentlicht: (2019)
von: Zhao, Pu, et al.
Veröffentlicht: (2019)
Red Teaming GPT-4V: Are GPT-4V Safe Against Uni/Multi-Modal Jailbreak Attacks?
von: Chen, Shuo, et al.
Veröffentlicht: (2024)
von: Chen, Shuo, et al.
Veröffentlicht: (2024)
MARAGE: Transferable Multi-Model Adversarial Attack for Retrieval-Augmented Generation Data Extraction
von: Hu, Xiao, et al.
Veröffentlicht: (2025)
von: Hu, Xiao, et al.
Veröffentlicht: (2025)
What Does the Server See? Understanding Privacy Leakage from Large Language Models in Split Inference
von: Fan, Mingyuan, et al.
Veröffentlicht: (2026)
von: Fan, Mingyuan, et al.
Veröffentlicht: (2026)
Proving membership in LLM pretraining data via data watermarks
von: Wei, Johnny Tian-Zheng, et al.
Veröffentlicht: (2024)
von: Wei, Johnny Tian-Zheng, et al.
Veröffentlicht: (2024)
Blind Baselines Beat Membership Inference Attacks for Foundation Models
von: Das, Debeshee, et al.
Veröffentlicht: (2024)
von: Das, Debeshee, et al.
Veröffentlicht: (2024)
LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet
von: Li, Nathaniel, et al.
Veröffentlicht: (2024)
von: Li, Nathaniel, et al.
Veröffentlicht: (2024)
On Calibration of LLM-based Guard Models for Reliable Content Moderation
von: Liu, Hongfu, et al.
Veröffentlicht: (2024)
von: Liu, Hongfu, et al.
Veröffentlicht: (2024)
Steering Dialogue Dynamics for Robustness against Multi-turn Jailbreaking Attacks
von: Hu, Hanjiang, et al.
Veröffentlicht: (2025)
von: Hu, Hanjiang, et al.
Veröffentlicht: (2025)
Attack and defense techniques in large language models: A survey and new perspectives
von: Liao, Zhiyu, et al.
Veröffentlicht: (2025)
von: Liao, Zhiyu, et al.
Veröffentlicht: (2025)
Composite Backdoor Attacks Against Large Language Models
von: Huang, Hai, et al.
Veröffentlicht: (2023)
von: Huang, Hai, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Synthetic Data Can Mislead Evaluations: Membership Inference as Machine Text Detection
von: Naseh, Ali, et al.
Veröffentlicht: (2025) -
MIA-Tuner: Adapting Large Language Models as Pre-training Text Detector
von: Fu, Wenjie, et al.
Veröffentlicht: (2024) -
Humanizing Machine-Generated Content: Evading AI-Text Detection through Adversarial Attack
von: Zhou, Ying, et al.
Veröffentlicht: (2024) -
AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
von: Zeng, Yifan, et al.
Veröffentlicht: (2024) -
GCG Attack On A Diffusion LLM
von: Neyroud, Ruben, et al.
Veröffentlicht: (2025)