Adversarial ML Problems Are Getting Harder to Solve and to Evaluate
Fuente:
arXiv
Guardado en:
| Autores principales: | Rando, Javier, Zhang, Jie, Carlini, Nicholas, Tramèr, Florian |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Adversarial Perturbations Cannot Reliably Protect Artists From Generative AI
por: Hönig, Robert, et al.
Publicado: (2024)
por: Hönig, Robert, et al.
Publicado: (2024)
AutoAdvExBench: Benchmarking autonomous exploitation of adversarial example defenses
por: Carlini, Nicholas, et al.
Publicado: (2025)
por: Carlini, Nicholas, et al.
Publicado: (2025)
Evading Black-box Classifiers Without Breaking Eggs
por: Debenedetti, Edoardo, et al.
Publicado: (2023)
por: Debenedetti, Edoardo, et al.
Publicado: (2023)
Position: Considerations for Differentially Private Learning with Large-Scale Public Pretraining
por: Tramèr, Florian, et al.
Publicado: (2022)
por: Tramèr, Florian, et al.
Publicado: (2022)
Universal Jailbreak Backdoors from Poisoned Human Feedback
por: Rando, Javier, et al.
Publicado: (2023)
por: Rando, Javier, et al.
Publicado: (2023)
Evaluating the Robustness of the "Ensemble Everything Everywhere" Defense
por: Zhang, Jie, et al.
Publicado: (2024)
por: Zhang, Jie, et al.
Publicado: (2024)
Laundering AI Authority with Adversarial Examples
por: Zhang, Jie, et al.
Publicado: (2026)
por: Zhang, Jie, et al.
Publicado: (2026)
Evaluations of Machine Learning Privacy Defenses are Misleading
por: Aerni, Michael, et al.
Publicado: (2024)
por: Aerni, Michael, et al.
Publicado: (2024)
Query-Based Adversarial Prompt Generation
por: Hayase, Jonathan, et al.
Publicado: (2024)
por: Hayase, Jonathan, et al.
Publicado: (2024)
Adversarial Search Engine Optimization for Large Language Models
por: Nestaas, Fredrik, et al.
Publicado: (2024)
por: Nestaas, Fredrik, et al.
Publicado: (2024)
An Adversarial Perspective on Machine Unlearning for AI Safety
por: Łucki, Jakub, et al.
Publicado: (2024)
por: Łucki, Jakub, et al.
Publicado: (2024)
Privacy Backdoors: Stealing Data with Corrupted Pretrained Models
por: Feng, Shanglun, et al.
Publicado: (2024)
por: Feng, Shanglun, et al.
Publicado: (2024)
Membership Inference Attacks on Sequence Models
por: Rossi, Lorenzo, et al.
Publicado: (2025)
por: Rossi, Lorenzo, et al.
Publicado: (2025)
Membership Inference Attacks Cannot Prove that a Model Was Trained On Your Data
por: Zhang, Jie, et al.
Publicado: (2024)
por: Zhang, Jie, et al.
Publicado: (2024)
Large-scale online deanonymization with LLMs
por: Lermen, Simon, et al.
Publicado: (2026)
por: Lermen, Simon, et al.
Publicado: (2026)
Blind Baselines Beat Membership Inference Attacks for Foundation Models
por: Das, Debeshee, et al.
Publicado: (2024)
por: Das, Debeshee, et al.
Publicado: (2024)
Exploring and Mitigating Adversarial Manipulation of Voting-Based Leaderboards
por: Huang, Yangsibo, et al.
Publicado: (2025)
por: Huang, Yangsibo, et al.
Publicado: (2025)
AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
por: Debenedetti, Edoardo, et al.
Publicado: (2024)
por: Debenedetti, Edoardo, et al.
Publicado: (2024)
Privacy Side Channels in Machine Learning Systems
por: Debenedetti, Edoardo, et al.
Publicado: (2023)
por: Debenedetti, Edoardo, et al.
Publicado: (2023)
Cutting through buggy adversarial example defenses: fixing 1 line of code breaks Sabre
por: Carlini, Nicholas
Publicado: (2024)
por: Carlini, Nicholas
Publicado: (2024)
Black-box Optimization of LLM Outputs by Asking for Directions
por: Zhang, Jie, et al.
Publicado: (2025)
por: Zhang, Jie, et al.
Publicado: (2025)
Poisoning Web-Scale Training Datasets is Practical
por: Carlini, Nicholas, et al.
Publicado: (2023)
por: Carlini, Nicholas, et al.
Publicado: (2023)
The Jailbreak Tax: How Useful are Your Jailbreak Outputs?
por: Nikolić, Kristina, et al.
Publicado: (2025)
por: Nikolić, Kristina, et al.
Publicado: (2025)
Whisper Smarter, not Harder: Adversarial Attack on Partial Suppression
por: Wong, Zheng Jie, et al.
Publicado: (2025)
por: Wong, Zheng Jie, et al.
Publicado: (2025)
Intriguing Properties of Adversarial ML Attacks in the Problem Space [Extended Version]
por: Cortellazzi, Jacopo, et al.
Publicado: (2019)
por: Cortellazzi, Jacopo, et al.
Publicado: (2019)
Remote Timing Attacks on Efficient Language Model Inference
por: Carlini, Nicholas, et al.
Publicado: (2024)
por: Carlini, Nicholas, et al.
Publicado: (2024)
Competition Report: Finding Universal Jailbreak Backdoors in Aligned LLMs
por: Rando, Javier, et al.
Publicado: (2024)
por: Rando, Javier, et al.
Publicado: (2024)
Evaluating the Robustness of a Production Malware Detection System to Transferable Adversarial Attacks
por: Nasr, Milad, et al.
Publicado: (2025)
por: Nasr, Milad, et al.
Publicado: (2025)
Persistent Pre-Training Poisoning of LLMs
por: Zhang, Yiming, et al.
Publicado: (2024)
por: Zhang, Yiming, et al.
Publicado: (2024)
The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
por: Nasr, Milad, et al.
Publicado: (2025)
por: Nasr, Milad, et al.
Publicado: (2025)
IF-GUIDE: Influence Function-Guided Detoxification of LLMs
por: Coalson, Zachary, et al.
Publicado: (2025)
por: Coalson, Zachary, et al.
Publicado: (2025)
SoK: Analyzing Adversarial Examples: A Framework to Study Adversary Knowledge
por: Fenaux, Lucas, et al.
Publicado: (2024)
por: Fenaux, Lucas, et al.
Publicado: (2024)
Development of an Edge Resilient ML Ensemble to Tolerate ICS Adversarial Attacks
por: Yao, Likai, et al.
Publicado: (2024)
por: Yao, Likai, et al.
Publicado: (2024)
Towards Sustainable SecureML: Quantifying Carbon Footprint of Adversarial Machine Learning
por: Hasan, Syed Mhamudul, et al.
Publicado: (2024)
por: Hasan, Syed Mhamudul, et al.
Publicado: (2024)
Gradient-based Jailbreak Images for Multimodal Fusion Models
por: Rando, Javier, et al.
Publicado: (2024)
por: Rando, Javier, et al.
Publicado: (2024)
A No-Defense Defense Against Gradient-Based Adversarial Attacks on ML-NIDS: Is Less More?
por: elShehaby, Mohamed, et al.
Publicado: (2026)
por: elShehaby, Mohamed, et al.
Publicado: (2026)
Privacy Backdoors: Enhancing Membership Inference through Poisoning Pre-trained Models
por: Wen, Yuxin, et al.
Publicado: (2024)
por: Wen, Yuxin, et al.
Publicado: (2024)
SoK: Watermarking for AI-Generated Content
por: Zhao, Xuandong, et al.
Publicado: (2024)
por: Zhao, Xuandong, et al.
Publicado: (2024)
Taking off the Rose-Tinted Glasses: A Critical Look at Adversarial ML Through the Lens of Evasion Attacks
por: Eykholt, Kevin, et al.
Publicado: (2024)
por: Eykholt, Kevin, et al.
Publicado: (2024)
SpinML: Customized Synthetic Data Generation for Private Training of Specialized ML Models
por: Zhang, Jiang, et al.
Publicado: (2025)
por: Zhang, Jiang, et al.
Publicado: (2025)
Ejemplares similares
-
Adversarial Perturbations Cannot Reliably Protect Artists From Generative AI
por: Hönig, Robert, et al.
Publicado: (2024) -
AutoAdvExBench: Benchmarking autonomous exploitation of adversarial example defenses
por: Carlini, Nicholas, et al.
Publicado: (2025) -
Evading Black-box Classifiers Without Breaking Eggs
por: Debenedetti, Edoardo, et al.
Publicado: (2023) -
Position: Considerations for Differentially Private Learning with Large-Scale Public Pretraining
por: Tramèr, Florian, et al.
Publicado: (2022) -
Universal Jailbreak Backdoors from Poisoned Human Feedback
por: Rando, Javier, et al.
Publicado: (2023)