Reformulation is All You Need: Addressing Malicious Text Features in DNNs
Fuente:
arXiv
Saved in:
| Main Authors: | Jiang, Yi, Ma, Oubo, Yang, Yong, Zhang, Tong, Ji, Shouling |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TooBadRL: Trigger Optimization to Boost Effectiveness of Backdoor Attacks on Deep Reinforcement Learning
by: Zhang, Mingxuan, et al.
Published: (2025)
by: Zhang, Mingxuan, et al.
Published: (2025)
Intrinsic Fingerprint of LLMs: Continue Training is NOT All You Need to Steal A Model!
by: Yoon, Do-hyeon, et al.
Published: (2025)
by: Yoon, Do-hyeon, et al.
Published: (2025)
SecEncoder: Logs are All You Need in Security
by: Bulut, Muhammed Fatih, et al.
Published: (2024)
by: Bulut, Muhammed Fatih, et al.
Published: (2024)
PRSA: Prompt Stealing Attacks against Real-World Prompt Services
by: Yang, Yong, et al.
Published: (2024)
by: Yang, Yong, et al.
Published: (2024)
Watermark under Fire: A Robustness Evaluation of LLM Watermarking
by: Liang, Jiacheng, et al.
Published: (2024)
by: Liang, Jiacheng, et al.
Published: (2024)
CLIBE: Detecting Dynamic Backdoors in Transformer-based NLP Models
by: Zeng, Rui, et al.
Published: (2024)
by: Zeng, Rui, et al.
Published: (2024)
Localizing Malicious Outputs from CodeLLM
by: Borana, Mayukh, et al.
Published: (2025)
by: Borana, Mayukh, et al.
Published: (2025)
All You Need is "Leet": Evading Hate-speech Detection AI
by: Kahu, Sampanna Yashwant, et al.
Published: (2025)
by: Kahu, Sampanna Yashwant, et al.
Published: (2025)
Angel or Demon: Investigating the Plasticity Interventions' Impact on Backdoor Threats in Deep Reinforcement Learning
by: Ma, Oubo, et al.
Published: (2026)
by: Ma, Oubo, et al.
Published: (2026)
UNIDOOR: A Universal Framework for Action-Level Backdoor Attacks in Deep Reinforcement Learning
by: Ma, Oubo, et al.
Published: (2025)
by: Ma, Oubo, et al.
Published: (2025)
MIA-Tuner: Adapting Large Language Models as Pre-training Text Detector
by: Fu, Wenjie, et al.
Published: (2024)
by: Fu, Wenjie, et al.
Published: (2024)
AEIOU: A Unified Defense Framework against NSFW Prompts in Text-to-Image Models
by: Wang, Yiming, et al.
Published: (2024)
by: Wang, Yiming, et al.
Published: (2024)
SUB-PLAY: Adversarial Policies against Partially Observed Multi-Agent Reinforcement Learning Systems
by: Ma, Oubo, et al.
Published: (2024)
by: Ma, Oubo, et al.
Published: (2024)
Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs
by: Cai, Will, et al.
Published: (2025)
by: Cai, Will, et al.
Published: (2025)
Covert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation
by: Halawi, Danny, et al.
Published: (2024)
by: Halawi, Danny, et al.
Published: (2024)
Watermarking Needs Input Repetition Masking
by: Khachaturov, David, et al.
Published: (2025)
by: Khachaturov, David, et al.
Published: (2025)
Transferable Embedding Inversion Attack: Uncovering Privacy Risks in Text Embeddings without Model Queries
by: Huang, Yu-Hsiang, et al.
Published: (2024)
by: Huang, Yu-Hsiang, et al.
Published: (2024)
If You Don't Understand It, Don't Use It: Eliminating Trojans with Filters Between Layers
by: Hernandez, Adriano
Published: (2024)
by: Hernandez, Adriano
Published: (2024)
FreqMark: Frequency-Based Watermark for Sentence-Level Detection of LLM-Generated Text
by: Xu, Zhenyu, et al.
Published: (2024)
by: Xu, Zhenyu, et al.
Published: (2024)
MOCHA: Are Code Language Models Robust Against Multi-Turn Malicious Coding Prompts?
by: Wahed, Muntasir, et al.
Published: (2025)
by: Wahed, Muntasir, et al.
Published: (2025)
Graded Suspiciousness of Adversarial Texts to Human
by: Tonni, Shakila Mahjabin, et al.
Published: (2024)
by: Tonni, Shakila Mahjabin, et al.
Published: (2024)
CR-UTP: Certified Robustness against Universal Text Perturbations on Large Language Models
by: Lou, Qian, et al.
Published: (2024)
by: Lou, Qian, et al.
Published: (2024)
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users
by: Hilel, Almog, et al.
Published: (2025)
by: Hilel, Almog, et al.
Published: (2025)
Text-CRS: A Generalized Certified Robustness Framework against Textual Adversarial Attacks
by: Zhang, Xinyu, et al.
Published: (2023)
by: Zhang, Xinyu, et al.
Published: (2023)
Practical Membership Inference Attacks against Fine-tuned Large Language Models via Self-prompt Calibration
by: Fu, Wenjie, et al.
Published: (2023)
by: Fu, Wenjie, et al.
Published: (2023)
"No Matter What You Do": Purifying GNN Models via Backdoor Unlearning
by: Zhang, Jiale, et al.
Published: (2024)
by: Zhang, Jiale, et al.
Published: (2024)
Malicious and Unintentional Disclosure Risks in Large Language Models for Code Generation
by: Rabin, Rafiqul, et al.
Published: (2025)
by: Rabin, Rafiqul, et al.
Published: (2025)
Differentially Private Knowledge Distillation via Synthetic Text Generation
by: Flemings, James, et al.
Published: (2024)
by: Flemings, James, et al.
Published: (2024)
CERT-ED: Certifiably Robust Text Classification for Edit Distance
by: Huang, Zhuoqun, et al.
Published: (2024)
by: Huang, Zhuoqun, et al.
Published: (2024)
Hidden in Plain Text: Emergence & Mitigation of Steganographic Collusion in LLMs
by: Mathew, Yohan, et al.
Published: (2024)
by: Mathew, Yohan, et al.
Published: (2024)
Semantic Stealth: Adversarial Text Attacks on NLP Using Several Methods
by: Dey, Roopkatha, et al.
Published: (2024)
by: Dey, Roopkatha, et al.
Published: (2024)
Differentially Private Synthetic Text Generation for Retrieval-Augmented Generation (RAG)
by: Mori, Junki, et al.
Published: (2025)
by: Mori, Junki, et al.
Published: (2025)
ContinuousBench: Can Differentially Private Synthetic Text Improve Capabilities?
by: Liu, Peihan, et al.
Published: (2026)
by: Liu, Peihan, et al.
Published: (2026)
Enhancing Adversarial Text Attacks on BERT Models with Projected Gradient Descent
by: Waghela, Hetvi, et al.
Published: (2024)
by: Waghela, Hetvi, et al.
Published: (2024)
The Canary's Echo: Auditing Privacy Risks of LLM-Generated Synthetic Text
by: Meeus, Matthieu, et al.
Published: (2025)
by: Meeus, Matthieu, et al.
Published: (2025)
TextSeal: A Localized LLM Watermark for Provenance & Distillation Protection
by: Sander, Tom, et al.
Published: (2026)
by: Sander, Tom, et al.
Published: (2026)
Assessing Deanonymization Risks with Stylometry-Assisted LLM Agent
by: Zhang, Boyang, et al.
Published: (2026)
by: Zhang, Boyang, et al.
Published: (2026)
A Modified Word Saliency-Based Adversarial Attack on Text Classification Models
by: Waghela, Hetvi, et al.
Published: (2024)
by: Waghela, Hetvi, et al.
Published: (2024)
Synthetic Data Can Mislead Evaluations: Membership Inference as Machine Text Detection
by: Naseh, Ali, et al.
Published: (2025)
by: Naseh, Ali, et al.
Published: (2025)
InvisibleInk: High-Utility and Low-Cost Text Generation with Differential Privacy
by: Vinod, Vishnu, et al.
Published: (2025)
by: Vinod, Vishnu, et al.
Published: (2025)
Similar Items
-
TooBadRL: Trigger Optimization to Boost Effectiveness of Backdoor Attacks on Deep Reinforcement Learning
by: Zhang, Mingxuan, et al.
Published: (2025) -
Intrinsic Fingerprint of LLMs: Continue Training is NOT All You Need to Steal A Model!
by: Yoon, Do-hyeon, et al.
Published: (2025) -
SecEncoder: Logs are All You Need in Security
by: Bulut, Muhammed Fatih, et al.
Published: (2024) -
PRSA: Prompt Stealing Attacks against Real-World Prompt Services
by: Yang, Yong, et al.
Published: (2024) -
Watermark under Fire: A Robustness Evaluation of LLM Watermarking
by: Liang, Jiacheng, et al.
Published: (2024)