Broken-Token: Filtering Obfuscated Prompts by Counting Characters-Per-Token
Fuente:
arXiv
Saved in:
| Main Authors: | Zychlinski, Shaked, Kainan, Yuval |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Whole New World: Creating a Parallel-Poisoned Web Only AI-Agents Can See
by: Zychlinski, Shaked
Published: (2025)
by: Zychlinski, Shaked
Published: (2025)
Do Stop Me Now: Detecting Boilerplate Responses with a Single Iteration
by: Kainan, Yuval, et al.
Published: (2025)
by: Kainan, Yuval, et al.
Published: (2025)
EmojiPrompt: Generative Prompt Obfuscation for Privacy-Preserving Communication with Cloud-based LLMs
by: Lin, Sam, et al.
Published: (2024)
by: Lin, Sam, et al.
Published: (2024)
Token-wise Influential Training Data Retrieval for Large Language Models
by: Lin, Huawei, et al.
Published: (2024)
by: Lin, Huawei, et al.
Published: (2024)
HeavyWater and SimplexWater: Distortion-Free LLM Watermarks for Low-Entropy Next-Token Predictions
by: Tsur, Dor, et al.
Published: (2025)
by: Tsur, Dor, et al.
Published: (2025)
Can LLMs Obfuscate Code? A Systematic Analysis of Large Language Models into Assembly Code Obfuscation
by: Mohseni, Seyedreza, et al.
Published: (2024)
by: Mohseni, Seyedreza, et al.
Published: (2024)
Auditing Pay-Per-Token in Large Language Models
by: Velasco, Ander Artola, et al.
Published: (2025)
by: Velasco, Ander Artola, et al.
Published: (2025)
Token-level Data Selection for Safe LLM Fine-tuning
by: Li, Yanping, et al.
Published: (2026)
by: Li, Yanping, et al.
Published: (2026)
Mask-GCG: Are All Tokens in Adversarial Suffixes Necessary for Jailbreak Attacks?
by: Mu, Junjie, et al.
Published: (2025)
by: Mu, Junjie, et al.
Published: (2025)
Backdoor Token Unlearning: Exposing and Defending Backdoors in Pretrained Language Models
by: Jiang, Peihai, et al.
Published: (2025)
by: Jiang, Peihai, et al.
Published: (2025)
Every Character Counts: From Vulnerability to Defense in Phishing Detection
by: Chiper, Maria, et al.
Published: (2025)
by: Chiper, Maria, et al.
Published: (2025)
Direct Token Optimization: A Self-contained Approach to Large Language Model Unlearning
by: Lee, Hong kyu, et al.
Published: (2025)
by: Lee, Hong kyu, et al.
Published: (2025)
STShield: Single-Token Sentinel for Real-Time Jailbreak Detection in Large Language Models
by: Wang, Xunguang, et al.
Published: (2025)
by: Wang, Xunguang, et al.
Published: (2025)
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense
by: Ouyang, Yang, et al.
Published: (2025)
by: Ouyang, Yang, et al.
Published: (2025)
Token Inflation: How Dishonest Providers Can Overcharge for Large Language Model Usage
by: Hoque, Shahinul, et al.
Published: (2026)
by: Hoque, Shahinul, et al.
Published: (2026)
Large Multimodal Agents for Accurate Phishing Detection with Enhanced Token Optimization and Cost Reduction
by: Trad, Fouad, et al.
Published: (2024)
by: Trad, Fouad, et al.
Published: (2024)
Forgetting-MarI: LLM Unlearning via Marginal Information Regularization
by: Xu, Shizhou, et al.
Published: (2025)
by: Xu, Shizhou, et al.
Published: (2025)
OD-Stega: LLM-Based Relatively Secure Steganography via Optimized Distributions
by: Huang, Yu-Shin, et al.
Published: (2024)
by: Huang, Yu-Shin, et al.
Published: (2024)
A Decision-Theoretic Formalisation of Steganography With Applications to LLM Monitoring
by: Anwar, Usman, et al.
Published: (2026)
by: Anwar, Usman, et al.
Published: (2026)
Protecting Privacy in Classifiers by Token Manipulation
by: Harel, Re'em, et al.
Published: (2024)
by: Harel, Re'em, et al.
Published: (2024)
Token-Level Privacy in Large Language Models
by: Harel, Re'em, et al.
Published: (2025)
by: Harel, Re'em, et al.
Published: (2025)
Interpreting the Repeated Token Phenomenon in Large Language Models
by: Yona, Itay, et al.
Published: (2025)
by: Yona, Itay, et al.
Published: (2025)
Sparse Tokens Suffice: Jailbreaking Audio Language Models via Token-Aware Gradient Optimization
by: Fang, Zheng, et al.
Published: (2026)
by: Fang, Zheng, et al.
Published: (2026)
SeqAR: Jailbreak LLMs with Sequential Auto-Generated Characters
by: Yang, Yan, et al.
Published: (2024)
by: Yang, Yan, et al.
Published: (2024)
A Generative Approach to LLM Harmfulness Mitigation with Red Flag Tokens
by: Dobre, David, et al.
Published: (2025)
by: Dobre, David, et al.
Published: (2025)
Token-Efficient Prompt Injection Attack: Provoking Cessation in LLM Reasoning via Adaptive Token Compression
by: Cui, Yu, et al.
Published: (2025)
by: Cui, Yu, et al.
Published: (2025)
How Different Tokenization Algorithms Impact LLMs and Transformer Models for Binary Code Analysis
by: Mostafa, Ahmed, et al.
Published: (2025)
by: Mostafa, Ahmed, et al.
Published: (2025)
DePrompt: Desensitization and Evaluation of Personal Identifiable Information in Large Language Model Prompts
by: Sun, Xiongtao, et al.
Published: (2024)
by: Sun, Xiongtao, et al.
Published: (2024)
Character as a Latent Variable in Large Language Models: A Mechanistic Account of Emergent Misalignment and Conditional Safety Failures
by: Su, Yanghao, et al.
Published: (2026)
by: Su, Yanghao, et al.
Published: (2026)
NeuroFilter: Privacy Guardrails for Conversational LLM Agents
by: Das, Saswat, et al.
Published: (2026)
by: Das, Saswat, et al.
Published: (2026)
Prompt Injection as Role Confusion
by: Ye, Charles, et al.
Published: (2026)
by: Ye, Charles, et al.
Published: (2026)
ObfuscaTune: Obfuscated Offsite Fine-tuning and Inference of Proprietary LLMs on Private Datasets
by: Frikha, Ahmed, et al.
Published: (2024)
by: Frikha, Ahmed, et al.
Published: (2024)
Soft Begging: Modular and Efficient Shielding of LLMs against Prompt Injection and Jailbreaking based on Prompt Tuning
by: Ostermann, Simon, et al.
Published: (2024)
by: Ostermann, Simon, et al.
Published: (2024)
PII-Bench: Evaluating Query-Aware Privacy Protection Systems
by: Shen, Hao, et al.
Published: (2025)
by: Shen, Hao, et al.
Published: (2025)
TITAN: Graph-Executable Reasoning for Cyber Threat Intelligence
by: Simoni, Marco, et al.
Published: (2025)
by: Simoni, Marco, et al.
Published: (2025)
DiscourseFlip: An Oblique Discourse-Level Opinion Manipulation Attack against Black-box Retrieval-Augmented Generation
by: Gong, Yuyang, et al.
Published: (2026)
by: Gong, Yuyang, et al.
Published: (2026)
Improving User Privacy in Personalized Generation: Client-Side Retrieval-Augmented Modification of Server-Side Generated Speculations
by: Salemi, Alireza, et al.
Published: (2026)
by: Salemi, Alireza, et al.
Published: (2026)
One Single Hub Text Breaks CLIP: Identifying Vulnerabilities in Cross-Modal Encoders via Hubness
by: Deguchi, Hiroyuki, et al.
Published: (2026)
by: Deguchi, Hiroyuki, et al.
Published: (2026)
SoK: Agentic Retrieval-Augmented Generation (RAG): Taxonomy, Architectures, Evaluation, and Research Directions
by: Mishra, Saroj, et al.
Published: (2026)
by: Mishra, Saroj, et al.
Published: (2026)
Are All Prompt Components Value-Neutral? Understanding the Heterogeneous Adversarial Robustness of Dissected Prompt in Large Language Models
by: Zheng, Yujia, et al.
Published: (2025)
by: Zheng, Yujia, et al.
Published: (2025)
Similar Items
-
A Whole New World: Creating a Parallel-Poisoned Web Only AI-Agents Can See
by: Zychlinski, Shaked
Published: (2025) -
Do Stop Me Now: Detecting Boilerplate Responses with a Single Iteration
by: Kainan, Yuval, et al.
Published: (2025) -
EmojiPrompt: Generative Prompt Obfuscation for Privacy-Preserving Communication with Cloud-based LLMs
by: Lin, Sam, et al.
Published: (2024) -
Token-wise Influential Training Data Retrieval for Large Language Models
by: Lin, Huawei, et al.
Published: (2024) -
HeavyWater and SimplexWater: Distortion-Free LLM Watermarks for Low-Entropy Next-Token Predictions
by: Tsur, Dor, et al.
Published: (2025)