Large-scale online deanonymization with LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Lermen, Simon, Paleka, Daniel, Swanson, Joshua, Aerni, Michael, Carlini, Nicholas, Tramèr, Florian |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluations of Machine Learning Privacy Defenses are Misleading
by: Aerni, Michael, et al.
Published: (2024)
by: Aerni, Michael, et al.
Published: (2024)
AutoAdvExBench: Benchmarking autonomous exploitation of adversarial example defenses
by: Carlini, Nicholas, et al.
Published: (2025)
by: Carlini, Nicholas, et al.
Published: (2025)
Position: Considerations for Differentially Private Learning with Large-Scale Public Pretraining
by: Tramèr, Florian, et al.
Published: (2022)
by: Tramèr, Florian, et al.
Published: (2022)
Membership Inference Attacks on Sequence Models
by: Rossi, Lorenzo, et al.
Published: (2025)
by: Rossi, Lorenzo, et al.
Published: (2025)
Query-Based Adversarial Prompt Generation
by: Hayase, Jonathan, et al.
Published: (2024)
by: Hayase, Jonathan, et al.
Published: (2024)
Evading Black-box Classifiers Without Breaking Eggs
by: Debenedetti, Edoardo, et al.
Published: (2023)
by: Debenedetti, Edoardo, et al.
Published: (2023)
Modal Aphasia: Can Unified Multimodal Models Describe Images From Memory?
by: Aerni, Michael, et al.
Published: (2025)
by: Aerni, Michael, et al.
Published: (2025)
Poisoning Web-Scale Training Datasets is Practical
by: Carlini, Nicholas, et al.
Published: (2023)
by: Carlini, Nicholas, et al.
Published: (2023)
Adversarial ML Problems Are Getting Harder to Solve and to Evaluate
by: Rando, Javier, et al.
Published: (2025)
by: Rando, Javier, et al.
Published: (2025)
Universal Jailbreak Backdoors from Poisoned Human Feedback
by: Rando, Javier, et al.
Published: (2023)
by: Rando, Javier, et al.
Published: (2023)
The Jailbreak Tax: How Useful are Your Jailbreak Outputs?
by: Nikolić, Kristina, et al.
Published: (2025)
by: Nikolić, Kristina, et al.
Published: (2025)
Persistent Pre-Training Poisoning of LLMs
by: Zhang, Yiming, et al.
Published: (2024)
by: Zhang, Yiming, et al.
Published: (2024)
SoK: Watermarking for AI-Generated Content
by: Zhao, Xuandong, et al.
Published: (2024)
by: Zhao, Xuandong, et al.
Published: (2024)
Competition Report: Finding Universal Jailbreak Backdoors in Aligned LLMs
by: Rando, Javier, et al.
Published: (2024)
by: Rando, Javier, et al.
Published: (2024)
LLMs unlock new paths to monetizing exploits
by: Carlini, Nicholas, et al.
Published: (2025)
by: Carlini, Nicholas, et al.
Published: (2025)
Are aligned neural networks adversarially aligned?
by: Carlini, Nicholas, et al.
Published: (2023)
by: Carlini, Nicholas, et al.
Published: (2023)
Evaluating the Robustness of the "Ensemble Everything Everywhere" Defense
by: Zhang, Jie, et al.
Published: (2024)
by: Zhang, Jie, et al.
Published: (2024)
An Adversarial Perspective on Machine Unlearning for AI Safety
by: Łucki, Jakub, et al.
Published: (2024)
by: Łucki, Jakub, et al.
Published: (2024)
Adversarial Search Engine Optimization for Large Language Models
by: Nestaas, Fredrik, et al.
Published: (2024)
by: Nestaas, Fredrik, et al.
Published: (2024)
Jailbreaking LLMs Without Gradients or Priors: Effective and Transferable Attacks
by: Nurlanov, Zhakshylyk, et al.
Published: (2026)
by: Nurlanov, Zhakshylyk, et al.
Published: (2026)
Stealing User Prompts from Mixture of Experts
by: Yona, Itay, et al.
Published: (2024)
by: Yona, Itay, et al.
Published: (2024)
Privacy Backdoors: Stealing Data with Corrupted Pretrained Models
by: Feng, Shanglun, et al.
Published: (2024)
by: Feng, Shanglun, et al.
Published: (2024)
Defeating Prompt Injections by Design
by: Debenedetti, Edoardo, et al.
Published: (2025)
by: Debenedetti, Edoardo, et al.
Published: (2025)
Permissioned LLMs: Enforcing Access Control in Large Language Models
by: Jayaraman, Bargav, et al.
Published: (2025)
by: Jayaraman, Bargav, et al.
Published: (2025)
Privacy Side Channels in Machine Learning Systems
by: Debenedetti, Edoardo, et al.
Published: (2023)
by: Debenedetti, Edoardo, et al.
Published: (2023)
Cutting through buggy adversarial example defenses: fixing 1 line of code breaks Sabre
by: Carlini, Nicholas
Published: (2024)
by: Carlini, Nicholas
Published: (2024)
Attacking LLMs and AI Agents: Advertisement Embedding Attacks Against Large Language Models
by: Guo, Qiming, et al.
Published: (2025)
by: Guo, Qiming, et al.
Published: (2025)
Can AI Models be Jailbroken to Phish Elderly Victims? An End-to-End Evaluation
by: Heiding, Fred, et al.
Published: (2025)
by: Heiding, Fred, et al.
Published: (2025)
Fast Exact Unlearning for In-Context Learning Data for LLMs
by: Muresanu, Andrei I., et al.
Published: (2024)
by: Muresanu, Andrei I., et al.
Published: (2024)
Sparse Models, Sparse Safety: Unsafe Routes in Mixture-of-Experts LLMs
by: Jiang, Yukun, et al.
Published: (2026)
by: Jiang, Yukun, et al.
Published: (2026)
ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?
by: Wang, Zhun, et al.
Published: (2026)
by: Wang, Zhun, et al.
Published: (2026)
IF-GUIDE: Influence Function-Guided Detoxification of LLMs
by: Coalson, Zachary, et al.
Published: (2025)
by: Coalson, Zachary, et al.
Published: (2025)
Automated Consistency Analysis of LLMs
by: Patwardhan, Aditya, et al.
Published: (2025)
by: Patwardhan, Aditya, et al.
Published: (2025)
AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs
by: Liu, Xiaogeng, et al.
Published: (2024)
by: Liu, Xiaogeng, et al.
Published: (2024)
Scaling Trends for Data Poisoning in LLMs
by: Bowen, Dillon, et al.
Published: (2024)
by: Bowen, Dillon, et al.
Published: (2024)
Can LLMs get help from other LLMs without revealing private information?
by: Hartmann, Florian, et al.
Published: (2024)
by: Hartmann, Florian, et al.
Published: (2024)
Trusted Machine Learning Models Unlock Private Inference for Problems Currently Infeasible with Cryptography
by: Shumailov, Ilia, et al.
Published: (2025)
by: Shumailov, Ilia, et al.
Published: (2025)
Leveraging RAG for Training-Free Alignment of LLMs
by: Halloran, John T.
Published: (2026)
by: Halloran, John T.
Published: (2026)
Injecting Universal Jailbreak Backdoors into LLMs in Minutes
by: Chen, Zhuowei, et al.
Published: (2025)
by: Chen, Zhuowei, et al.
Published: (2025)
DPrivBench: Benchmarking LLMs' Reasoning for Differential Privacy
by: Wang, Erchi, et al.
Published: (2026)
by: Wang, Erchi, et al.
Published: (2026)
Similar Items
-
Evaluations of Machine Learning Privacy Defenses are Misleading
by: Aerni, Michael, et al.
Published: (2024) -
AutoAdvExBench: Benchmarking autonomous exploitation of adversarial example defenses
by: Carlini, Nicholas, et al.
Published: (2025) -
Position: Considerations for Differentially Private Learning with Large-Scale Public Pretraining
by: Tramèr, Florian, et al.
Published: (2022) -
Membership Inference Attacks on Sequence Models
by: Rossi, Lorenzo, et al.
Published: (2025) -
Query-Based Adversarial Prompt Generation
by: Hayase, Jonathan, et al.
Published: (2024)