Evading Black-box Classifiers Without Breaking Eggs
Fuente:
arXiv
Guardado en:
| Autores principales: | Debenedetti, Edoardo, Carlini, Nicholas, Tramèr, Florian |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
AutoAdvExBench: Benchmarking autonomous exploitation of adversarial example defenses
por: Carlini, Nicholas, et al.
Publicado: (2025)
por: Carlini, Nicholas, et al.
Publicado: (2025)
Adversarial Search Engine Optimization for Large Language Models
por: Nestaas, Fredrik, et al.
Publicado: (2024)
por: Nestaas, Fredrik, et al.
Publicado: (2024)
Privacy Side Channels in Machine Learning Systems
por: Debenedetti, Edoardo, et al.
Publicado: (2023)
por: Debenedetti, Edoardo, et al.
Publicado: (2023)
Position: Considerations for Differentially Private Learning with Large-Scale Public Pretraining
por: Tramèr, Florian, et al.
Publicado: (2022)
por: Tramèr, Florian, et al.
Publicado: (2022)
AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
por: Debenedetti, Edoardo, et al.
Publicado: (2024)
por: Debenedetti, Edoardo, et al.
Publicado: (2024)
Adversarial ML Problems Are Getting Harder to Solve and to Evaluate
por: Rando, Javier, et al.
Publicado: (2025)
por: Rando, Javier, et al.
Publicado: (2025)
Black-box Optimization of LLM Outputs by Asking for Directions
por: Zhang, Jie, et al.
Publicado: (2025)
por: Zhang, Jie, et al.
Publicado: (2025)
EvadeDroid: A Practical Evasion Attack on Machine Learning for Black-box Android Malware Detection
por: Bostani, Hamid, et al.
Publicado: (2021)
por: Bostani, Hamid, et al.
Publicado: (2021)
Evaluating the Robustness of the "Ensemble Everything Everywhere" Defense
por: Zhang, Jie, et al.
Publicado: (2024)
por: Zhang, Jie, et al.
Publicado: (2024)
Privacy Backdoors: Stealing Data with Corrupted Pretrained Models
por: Feng, Shanglun, et al.
Publicado: (2024)
por: Feng, Shanglun, et al.
Publicado: (2024)
Large-scale online deanonymization with LLMs
por: Lermen, Simon, et al.
Publicado: (2026)
por: Lermen, Simon, et al.
Publicado: (2026)
Query-Based Adversarial Prompt Generation
por: Hayase, Jonathan, et al.
Publicado: (2024)
por: Hayase, Jonathan, et al.
Publicado: (2024)
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
por: Chao, Patrick, et al.
Publicado: (2024)
por: Chao, Patrick, et al.
Publicado: (2024)
Cutting through buggy adversarial example defenses: fixing 1 line of code breaks Sabre
por: Carlini, Nicholas
Publicado: (2024)
por: Carlini, Nicholas
Publicado: (2024)
Poisoning Web-Scale Training Datasets is Practical
por: Carlini, Nicholas, et al.
Publicado: (2023)
por: Carlini, Nicholas, et al.
Publicado: (2023)
Evaluations of Machine Learning Privacy Defenses are Misleading
por: Aerni, Michael, et al.
Publicado: (2024)
por: Aerni, Michael, et al.
Publicado: (2024)
LLMs unlock new paths to monetizing exploits
por: Carlini, Nicholas, et al.
Publicado: (2025)
por: Carlini, Nicholas, et al.
Publicado: (2025)
Design Patterns for Securing LLM Agents against Prompt Injections
por: Beurer-Kellner, Luca, et al.
Publicado: (2025)
por: Beurer-Kellner, Luca, et al.
Publicado: (2025)
Adversarial Perturbations Cannot Reliably Protect Artists From Generative AI
por: Hönig, Robert, et al.
Publicado: (2024)
por: Hönig, Robert, et al.
Publicado: (2024)
Membership Inference Attacks on Sequence Models
por: Rossi, Lorenzo, et al.
Publicado: (2025)
por: Rossi, Lorenzo, et al.
Publicado: (2025)
Membership Inference Attacks Cannot Prove that a Model Was Trained On Your Data
por: Zhang, Jie, et al.
Publicado: (2024)
por: Zhang, Jie, et al.
Publicado: (2024)
Laundering AI Authority with Adversarial Examples
por: Zhang, Jie, et al.
Publicado: (2026)
por: Zhang, Jie, et al.
Publicado: (2026)
Remote Timing Attacks on Efficient Language Model Inference
por: Carlini, Nicholas, et al.
Publicado: (2024)
por: Carlini, Nicholas, et al.
Publicado: (2024)
Defeating Prompt Injections by Design
por: Debenedetti, Edoardo, et al.
Publicado: (2025)
por: Debenedetti, Edoardo, et al.
Publicado: (2025)
Blind Baselines Beat Membership Inference Attacks for Foundation Models
por: Das, Debeshee, et al.
Publicado: (2024)
por: Das, Debeshee, et al.
Publicado: (2024)
Universal Jailbreak Backdoors from Poisoned Human Feedback
por: Rando, Javier, et al.
Publicado: (2023)
por: Rando, Javier, et al.
Publicado: (2023)
Traceable Black-box Watermarks for Federated Learning
por: Xu, Jiahao, et al.
Publicado: (2025)
por: Xu, Jiahao, et al.
Publicado: (2025)
The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
por: Nasr, Milad, et al.
Publicado: (2025)
por: Nasr, Milad, et al.
Publicado: (2025)
Exploring and Mitigating Adversarial Manipulation of Voting-Based Leaderboards
por: Huang, Yangsibo, et al.
Publicado: (2025)
por: Huang, Yangsibo, et al.
Publicado: (2025)
The Jailbreak Tax: How Useful are Your Jailbreak Outputs?
por: Nikolić, Kristina, et al.
Publicado: (2025)
por: Nikolić, Kristina, et al.
Publicado: (2025)
IF-GUIDE: Influence Function-Guided Detoxification of LLMs
por: Coalson, Zachary, et al.
Publicado: (2025)
por: Coalson, Zachary, et al.
Publicado: (2025)
Watch Out! Simple Horizontal Class Backdoor Can Trivially Evade Defense
por: Ma, Hua, et al.
Publicado: (2023)
por: Ma, Hua, et al.
Publicado: (2023)
Black-box Adversarial Transferability: An Empirical Study in Cybersecurity Perspective
por: Roshan, Khushnaseeb, et al.
Publicado: (2024)
por: Roshan, Khushnaseeb, et al.
Publicado: (2024)
Dynamic Black-box Backdoor Attacks on IoT Sensory Data
por: Chathoth, Ajesh Koyatan, et al.
Publicado: (2025)
por: Chathoth, Ajesh Koyatan, et al.
Publicado: (2025)
SEA: Shareable and Explainable Attribution for Query-based Black-box Attacks
por: Gao, Yue, et al.
Publicado: (2023)
por: Gao, Yue, et al.
Publicado: (2023)
LoRAGuard: An Effective Black-box Watermarking Approach for LoRAs
por: Lv, Peizhuo, et al.
Publicado: (2025)
por: Lv, Peizhuo, et al.
Publicado: (2025)
Online Poisoning Attack Against Reinforcement Learning under Black-box Environments
por: Li, Jianhui, et al.
Publicado: (2024)
por: Li, Jianhui, et al.
Publicado: (2024)
MergePrint: Merge-Resistant Fingerprints for Robust Black-box Ownership Verification of Large Language Models
por: Yamabe, Shojiro, et al.
Publicado: (2024)
por: Yamabe, Shojiro, et al.
Publicado: (2024)
Privacy Backdoors: Enhancing Membership Inference through Poisoning Pre-trained Models
por: Wen, Yuxin, et al.
Publicado: (2024)
por: Wen, Yuxin, et al.
Publicado: (2024)
AuthorMist: Evading AI Text Detectors with Reinforcement Learning
por: David, Isaac, et al.
Publicado: (2025)
por: David, Isaac, et al.
Publicado: (2025)
Ejemplares similares
-
AutoAdvExBench: Benchmarking autonomous exploitation of adversarial example defenses
por: Carlini, Nicholas, et al.
Publicado: (2025) -
Adversarial Search Engine Optimization for Large Language Models
por: Nestaas, Fredrik, et al.
Publicado: (2024) -
Privacy Side Channels in Machine Learning Systems
por: Debenedetti, Edoardo, et al.
Publicado: (2023) -
Position: Considerations for Differentially Private Learning with Large-Scale Public Pretraining
por: Tramèr, Florian, et al.
Publicado: (2022) -
AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
por: Debenedetti, Edoardo, et al.
Publicado: (2024)