On the Vulnerability of Text Sanitization
Fuente:
arXiv
Saved in:
| Main Authors: | Tong, Meng, Chen, Kejiang, Yuan, Xiaojian, Liu, Jiayang, Zhang, Weiming, Yu, Nenghai, Zhang, Jie |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Silent Guardian: Protecting Text from Malicious Exploitation by Large Language Models
by: Zhao, Jiawei, et al.
Published: (2023)
by: Zhao, Jiawei, et al.
Published: (2023)
SQL Injection Jailbreak: A Structural Disaster of Large Language Models
by: Zhao, Jiawei, et al.
Published: (2024)
by: Zhao, Jiawei, et al.
Published: (2024)
Performance-lossless Black-box Model Watermarking
by: Zhao, Na, et al.
Published: (2023)
by: Zhao, Na, et al.
Published: (2023)
InferDPT: Privacy-Preserving Inference for Closed-box Large Language Model
by: Tong, Meng, et al.
Published: (2023)
by: Tong, Meng, et al.
Published: (2023)
A high-capacity linguistic steganography based on entropy-driven rank-token mapping
by: Jiang, Jun, et al.
Published: (2025)
by: Jiang, Jun, et al.
Published: (2025)
PSRT: Accelerating LRM-based Guard Models via Prefilled Safe Reasoning Traces
by: Zhao, Jiawei, et al.
Published: (2025)
by: Zhao, Jiawei, et al.
Published: (2025)
GIFDL: Generated Image Fluctuation Distortion Learning for Enhancing Steganographic Security
by: Wang, Xiangkun, et al.
Published: (2025)
by: Wang, Xiangkun, et al.
Published: (2025)
Provably Secure Public-Key Steganography Based on Admissible Encoding
by: Zhang, Xin, et al.
Published: (2025)
by: Zhang, Xin, et al.
Published: (2025)
Turning Your Strength into Watermark: Watermarking Large Language Model via Knowledge Injection
by: Li, Shuai, et al.
Published: (2023)
by: Li, Shuai, et al.
Published: (2023)
Provably Secure Disambiguating Neural Linguistic Steganography
by: Qi, Yuang, et al.
Published: (2024)
by: Qi, Yuang, et al.
Published: (2024)
Provably Secure Agent Guardrail
by: Wu, Benlong, et al.
Published: (2026)
by: Wu, Benlong, et al.
Published: (2026)
Prefix Guidance: A Steering Wheel for Large Language Models to Defend Against Jailbreak Attacks
by: Zhao, Jiawei, et al.
Published: (2024)
by: Zhao, Jiawei, et al.
Published: (2024)
EditMark: Watermarking Large Language Models based on Model Editing
by: Li, Shuai, et al.
Published: (2025)
by: Li, Shuai, et al.
Published: (2025)
STEAD: Robust Provably Secure Linguistic Steganography with Diffusion Language Model
by: Qi, Yuang, et al.
Published: (2026)
by: Qi, Yuang, et al.
Published: (2026)
Natias: Neuron Attribution based Transferable Image Adversarial Steganography
by: Fan, Zexin, et al.
Published: (2024)
by: Fan, Zexin, et al.
Published: (2024)
De-AntiFake: Rethinking the Protective Perturbations Against Voice Cloning Attacks
by: Fan, Wei, et al.
Published: (2025)
by: Fan, Wei, et al.
Published: (2025)
Gaussian Shading: Provable Performance-Lossless Image Watermarking for Diffusion Models
by: Yang, Zijin, et al.
Published: (2024)
by: Yang, Zijin, et al.
Published: (2024)
Membership Inference Attacks on Tokenizers of Large Language Models
by: Tong, Meng, et al.
Published: (2025)
by: Tong, Meng, et al.
Published: (2025)
WavInWav: Time-domain Speech Hiding via Invertible Neural Network
by: Fan, Wei, et al.
Published: (2025)
by: Fan, Wei, et al.
Published: (2025)
SemBind: Binding Diffusion Watermarks to Semantics Against Black-Box Forgery Attacks
by: Zhang, Xin, et al.
Published: (2026)
by: Zhang, Xin, et al.
Published: (2026)
AutoPT: How Far Are We from the End2End Automated Web Penetration Testing?
by: Wu, Benlong, et al.
Published: (2024)
by: Wu, Benlong, et al.
Published: (2024)
LiteUpdate: A Lightweight Framework for Updating AI-Generated Image Detectors
by: Lu, Jiajie, et al.
Published: (2025)
by: Lu, Jiajie, et al.
Published: (2025)
Anota: Identifying Business Logic Vulnerabilities via Annotation-Based Sanitization
by: Wang, Meng, et al.
Published: (2025)
by: Wang, Meng, et al.
Published: (2025)
Gaussian Shading++: Rethinking the Realistic Deployment Challenge of Performance-Lossless Image Watermark for Diffusion Models
by: Yang, Zijin, et al.
Published: (2025)
by: Yang, Zijin, et al.
Published: (2025)
FoC: Figure out the Cryptographic Functions in Stripped Binaries with LLMs
by: Shang, Xiuwei, et al.
Published: (2024)
by: Shang, Xiuwei, et al.
Published: (2024)
AquaLoRA: Toward White-box Protection for Customized Stable Diffusion Models via Watermark LoRA
by: Feng, Weitao, et al.
Published: (2024)
by: Feng, Weitao, et al.
Published: (2024)
BURN: Backdoor Unlearning via Adversarial Boundary Analysis
by: Su, Yanghao, et al.
Published: (2025)
by: Su, Yanghao, et al.
Published: (2025)
Exploiting Vulnerabilities in Speech Translation Systems through Targeted Adversarial Attacks
by: Liu, Chang, et al.
Published: (2025)
by: Liu, Chang, et al.
Published: (2025)
Model X-ray:Detecting Backdoored Models via Decision Boundary
by: Su, Yanghao, et al.
Published: (2024)
by: Su, Yanghao, et al.
Published: (2024)
State-Dependent Safety Failures in Multi-Turn Language Model Interaction
by: Li, Pengcheng, et al.
Published: (2026)
by: Li, Pengcheng, et al.
Published: (2026)
©Plug-in Authorization for Human Content Copyright Protection in Text-to-Image Model
by: Zhou, Chao, et al.
Published: (2024)
by: Zhou, Chao, et al.
Published: (2024)
Beyond the Edge of Function: Unraveling the Patterns of Type Recovery in Binary Code
by: Li, Gangyang, et al.
Published: (2025)
by: Li, Gangyang, et al.
Published: (2025)
Demystifying RCE Vulnerabilities in LLM-Integrated Apps
by: Liu, Tong, et al.
Published: (2023)
by: Liu, Tong, et al.
Published: (2023)
Character as a Latent Variable in Large Language Models: A Mechanistic Account of Emergent Misalignment and Conditional Safety Failures
by: Su, Yanghao, et al.
Published: (2026)
by: Su, Yanghao, et al.
Published: (2026)
Safe Text-to-Image Generation: Simply Sanitize the Prompt Embedding
by: Qiu, Huming, et al.
Published: (2024)
by: Qiu, Huming, et al.
Published: (2024)
SparSamp: Efficient Provably Secure Steganography Based on Sparse Sampling
by: Wang, Yaofei, et al.
Published: (2025)
by: Wang, Yaofei, et al.
Published: (2025)
SiGRRW: A Single-Watermark Robust Reversible Watermarking Framework with Guiding Strategy
by: Xu, Zikai, et al.
Published: (2026)
by: Xu, Zikai, et al.
Published: (2026)
SafeGuider: Robust and Practical Content Safety Control for Text-to-Image Models
by: Qi, Peigui, et al.
Published: (2025)
by: Qi, Peigui, et al.
Published: (2025)
How Far Have We Gone in Binary Code Understanding Using Large Language Models
by: Shang, Xiuwei, et al.
Published: (2024)
by: Shang, Xiuwei, et al.
Published: (2024)
The Double-edged Sword of LLM-based Data Reconstruction: Understanding and Mitigating Contextual Vulnerability in Word-level Differential Privacy Text Sanitization
by: Meisenbacher, Stephen, et al.
Published: (2025)
by: Meisenbacher, Stephen, et al.
Published: (2025)
Similar Items
-
Silent Guardian: Protecting Text from Malicious Exploitation by Large Language Models
by: Zhao, Jiawei, et al.
Published: (2023) -
SQL Injection Jailbreak: A Structural Disaster of Large Language Models
by: Zhao, Jiawei, et al.
Published: (2024) -
Performance-lossless Black-box Model Watermarking
by: Zhao, Na, et al.
Published: (2023) -
InferDPT: Privacy-Preserving Inference for Closed-box Large Language Model
by: Tong, Meng, et al.
Published: (2023) -
A high-capacity linguistic steganography based on entropy-driven rank-token mapping
by: Jiang, Jun, et al.
Published: (2025)