AlienLM: Alienization of Language for API-Boundary Privacy in Black-Box LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Jaehee, Kang, Pilsung |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Jailbreaking Commercial Black-Box LLMs with Explicitly Harmful Prompts
by: Zhang, Chiyu, et al.
Published: (2025)
by: Zhang, Chiyu, et al.
Published: (2025)
PrivLM-Bench: A Multi-level Privacy Evaluation Benchmark for Language Models
by: Li, Haoran, et al.
Published: (2023)
by: Li, Haoran, et al.
Published: (2023)
EvoDefense: Co-Evolving Black-Box Defense with Large Language Models
by: Li, Yu, et al.
Published: (2026)
by: Li, Yu, et al.
Published: (2026)
FlexLLM: Exploring LLM Customization for Moving Target Defense on Black-Box LLMs Against Jailbreak Attacks
by: Chen, Bocheng, et al.
Published: (2024)
by: Chen, Bocheng, et al.
Published: (2024)
Black-Box Guardrail Reverse-engineering Attack
by: Yao, Hongwei, et al.
Published: (2025)
by: Yao, Hongwei, et al.
Published: (2025)
The Model's Language Matters: A Comparative Privacy Analysis of LLMs
by: Mishra, Abhishek K., et al.
Published: (2025)
by: Mishra, Abhishek K., et al.
Published: (2025)
Graph of Attacks: Improved Black-Box and Interpretable Jailbreaks for LLMs
by: Akbar-Tajari, Mohammad, et al.
Published: (2025)
by: Akbar-Tajari, Mohammad, et al.
Published: (2025)
A Watermark for Black-Box Language Models
by: Bahri, Dara, et al.
Published: (2024)
by: Bahri, Dara, et al.
Published: (2024)
Effective and Efficient Jailbreaks of Black-Box LLMs with Cross-Behavior Attacks
by: Gohil, Vasudev
Published: (2025)
by: Gohil, Vasudev
Published: (2025)
BinarySelect to Improve Accessibility of Black-Box Attack Research
by: Ghosh, Shatarupa, et al.
Published: (2024)
by: Ghosh, Shatarupa, et al.
Published: (2024)
Cross-Lingual Summarization as a Black-Box Watermark Removal Attack
by: Ganesan, Gokul
Published: (2025)
by: Ganesan, Gokul
Published: (2025)
PSM: Prompt Sensitivity Minimization via LLM-Guided Black-Box Optimization
by: Jawad, Huseein, et al.
Published: (2025)
by: Jawad, Huseein, et al.
Published: (2025)
Auto-Tuning Safety Guardrails for Black-Box Large Language Models
by: Abdulkadir, Perry
Published: (2025)
by: Abdulkadir, Perry
Published: (2025)
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
by: Mehrotra, Anay, et al.
Published: (2023)
by: Mehrotra, Anay, et al.
Published: (2023)
Privacy in Large Language Models: Attacks, Defenses and Future Directions
by: Li, Haoran, et al.
Published: (2023)
by: Li, Haoran, et al.
Published: (2023)
Continual Pretraining on Encrypted Synthetic Data for Privacy-Preserving LLMs
by: Liu, Honghao, et al.
Published: (2026)
by: Liu, Honghao, et al.
Published: (2026)
RTD-Guard: A Black-Box Textual Adversarial Detection Framework via Replacement Token Detection
by: Zhu, He, et al.
Published: (2026)
by: Zhu, He, et al.
Published: (2026)
Rethinking LLM Watermark Detection in Black-Box Settings: A Non-Intrusive Third-Party Framework
by: Wang, Zhuoshang, et al.
Published: (2026)
by: Wang, Zhuoshang, et al.
Published: (2026)
Black-Box Opinion Manipulation Attacks to Retrieval-Augmented Generation of Large Language Models
by: Chen, Zhuo, et al.
Published: (2024)
by: Chen, Zhuo, et al.
Published: (2024)
Token-Level Privacy in Large Language Models
by: Harel, Re'em, et al.
Published: (2025)
by: Harel, Re'em, et al.
Published: (2025)
PIG: Privacy Jailbreak Attack on LLMs via Gradient-based Iterative In-Context Optimization
by: Wang, Yidan, et al.
Published: (2025)
by: Wang, Yidan, et al.
Published: (2025)
Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory
by: Mireshghallah, Niloofar, et al.
Published: (2023)
by: Mireshghallah, Niloofar, et al.
Published: (2023)
Privacy-Preserving Instructions for Aligning Large Language Models
by: Yu, Da, et al.
Published: (2024)
by: Yu, Da, et al.
Published: (2024)
A General Pseudonymization Framework for Cloud-Based LLMs: Replacing Privacy Information in Controlled Text Generation
by: Hou, Shilong, et al.
Published: (2025)
by: Hou, Shilong, et al.
Published: (2025)
Your Inference Request Will Become a Black Box: Confidential Inference for Cloud-based Large Language Models
by: Huang, Chung-ju, et al.
Published: (2026)
by: Huang, Chung-ju, et al.
Published: (2026)
"Moralized" Multi-Step Jailbreak Prompts: Black-Box Testing of Guardrails in Large Language Models for Verbal Attacks
by: Wang, Libo
Published: (2024)
by: Wang, Libo
Published: (2024)
Break the Breakout: Reinventing LM Defense Against Jailbreak Attacks with Self-Refinement
by: Kim, Heegyu, et al.
Published: (2024)
by: Kim, Heegyu, et al.
Published: (2024)
PAPILLON: Privacy Preservation from Internet-based and Local Language Model Ensembles
by: Siyan, Li, et al.
Published: (2024)
by: Siyan, Li, et al.
Published: (2024)
Privacy-Preserving Parameter-Efficient Fine-Tuning for Large Language Model Services
by: Li, Yansong, et al.
Published: (2023)
by: Li, Yansong, et al.
Published: (2023)
PAL: Proxy-Guided Black-Box Attack on Large Language Models
by: Sitawarin, Chawin, et al.
Published: (2024)
by: Sitawarin, Chawin, et al.
Published: (2024)
TrailBlazer: History-Guided Reinforcement Learning for Black-Box LLM Jailbreaking
by: Yoon, Sung-Hoon, et al.
Published: (2026)
by: Yoon, Sung-Hoon, et al.
Published: (2026)
Auditing Black-Box LLM APIs with a Rank-Based Uniformity Test
by: Zhu, Xiaoyuan, et al.
Published: (2025)
by: Zhu, Xiaoyuan, et al.
Published: (2025)
Privacy Ripple Effects from Adding or Removing Personal Information in Language Model Training
by: Borkar, Jaydeep, et al.
Published: (2025)
by: Borkar, Jaydeep, et al.
Published: (2025)
Privacy Checklist: Privacy Violation Detection Grounding on Contextual Integrity Theory
by: Li, Haoran, et al.
Published: (2024)
by: Li, Haoran, et al.
Published: (2024)
TrojanStego: Your Language Model Can Secretly Be A Steganographic Privacy Leaking Agent
by: Meier, Dominik, et al.
Published: (2025)
by: Meier, Dominik, et al.
Published: (2025)
GoldCoin: Grounding Large Language Models in Privacy Laws via Contextual Integrity Theory
by: Fan, Wei, et al.
Published: (2024)
by: Fan, Wei, et al.
Published: (2024)
PATCH: Mitigating PII Leakage in Language Models with Privacy-Aware Targeted Circuit PatcHing
by: Hughes, Anthony, et al.
Published: (2025)
by: Hughes, Anthony, et al.
Published: (2025)
MemPrivacy: Privacy-Preserving Personalized Memory Management for Edge-Cloud Agents
by: Chen, Yining, et al.
Published: (2026)
by: Chen, Yining, et al.
Published: (2026)
Privacy-R1: Privacy-Aware Multi-LLM Agent Collaboration via Reinforcement Learning
by: Hui, Zheng, et al.
Published: (2025)
by: Hui, Zheng, et al.
Published: (2025)
Beyond Theoretical Bounds: Empirical Privacy Loss Calibration for Text Rewriting Under Local Differential Privacy
by: Li, Weijun, et al.
Published: (2026)
by: Li, Weijun, et al.
Published: (2026)
Similar Items
-
Jailbreaking Commercial Black-Box LLMs with Explicitly Harmful Prompts
by: Zhang, Chiyu, et al.
Published: (2025) -
PrivLM-Bench: A Multi-level Privacy Evaluation Benchmark for Language Models
by: Li, Haoran, et al.
Published: (2023) -
EvoDefense: Co-Evolving Black-Box Defense with Large Language Models
by: Li, Yu, et al.
Published: (2026) -
FlexLLM: Exploring LLM Customization for Moving Target Defense on Black-Box LLMs Against Jailbreak Attacks
by: Chen, Bocheng, et al.
Published: (2024) -
Black-Box Guardrail Reverse-engineering Attack
by: Yao, Hongwei, et al.
Published: (2025)