Laundering AI Authority with Adversarial Examples
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Jie, Peetathawatchai, Pura, Tramèr, Florian, Shafran, Avital |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Adversarial ML Problems Are Getting Harder to Solve and to Evaluate
di: Rando, Javier, et al.
Pubblicazione: (2025)
di: Rando, Javier, et al.
Pubblicazione: (2025)
Rerouting LLM Routers
di: Shafran, Avital, et al.
Pubblicazione: (2025)
di: Shafran, Avital, et al.
Pubblicazione: (2025)
Machine Against the RAG: Jamming Retrieval-Augmented Generation with Blocker Documents
di: Shafran, Avital, et al.
Pubblicazione: (2024)
di: Shafran, Avital, et al.
Pubblicazione: (2024)
Adversarial Search Engine Optimization for Large Language Models
di: Nestaas, Fredrik, et al.
Pubblicazione: (2024)
di: Nestaas, Fredrik, et al.
Pubblicazione: (2024)
Beyond Labeling Oracles: What does it mean to steal ML models?
di: Shafran, Avital, et al.
Pubblicazione: (2023)
di: Shafran, Avital, et al.
Pubblicazione: (2023)
Evaluations of Machine Learning Privacy Defenses are Misleading
di: Aerni, Michael, et al.
Pubblicazione: (2024)
di: Aerni, Michael, et al.
Pubblicazione: (2024)
Privacy Backdoors: Stealing Data with Corrupted Pretrained Models
di: Feng, Shanglun, et al.
Pubblicazione: (2024)
di: Feng, Shanglun, et al.
Pubblicazione: (2024)
Membership Inference Attacks on Sequence Models
di: Rossi, Lorenzo, et al.
Pubblicazione: (2025)
di: Rossi, Lorenzo, et al.
Pubblicazione: (2025)
Membership Inference Attacks Cannot Prove that a Model Was Trained On Your Data
di: Zhang, Jie, et al.
Pubblicazione: (2024)
di: Zhang, Jie, et al.
Pubblicazione: (2024)
Blind Baselines Beat Membership Inference Attacks for Foundation Models
di: Das, Debeshee, et al.
Pubblicazione: (2024)
di: Das, Debeshee, et al.
Pubblicazione: (2024)
SoK: Analyzing Adversarial Examples: A Framework to Study Adversary Knowledge
di: Fenaux, Lucas, et al.
Pubblicazione: (2024)
di: Fenaux, Lucas, et al.
Pubblicazione: (2024)
Black-box Optimization of LLM Outputs by Asking for Directions
di: Zhang, Jie, et al.
Pubblicazione: (2025)
di: Zhang, Jie, et al.
Pubblicazione: (2025)
Evading Black-box Classifiers Without Breaking Eggs
di: Debenedetti, Edoardo, et al.
Pubblicazione: (2023)
di: Debenedetti, Edoardo, et al.
Pubblicazione: (2023)
Position: Considerations for Differentially Private Learning with Large-Scale Public Pretraining
di: Tramèr, Florian, et al.
Pubblicazione: (2022)
di: Tramèr, Florian, et al.
Pubblicazione: (2022)
The Jailbreak Tax: How Useful are Your Jailbreak Outputs?
di: Nikolić, Kristina, et al.
Pubblicazione: (2025)
di: Nikolić, Kristina, et al.
Pubblicazione: (2025)
An Adversarial Perspective on Machine Unlearning for AI Safety
di: Łucki, Jakub, et al.
Pubblicazione: (2024)
di: Łucki, Jakub, et al.
Pubblicazione: (2024)
Detecting Adversarial Examples
di: Mumcu, Furkan, et al.
Pubblicazione: (2024)
di: Mumcu, Furkan, et al.
Pubblicazione: (2024)
AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
di: Debenedetti, Edoardo, et al.
Pubblicazione: (2024)
di: Debenedetti, Edoardo, et al.
Pubblicazione: (2024)
Transferability Ranking of Adversarial Examples
di: Levy, Mosh, et al.
Pubblicazione: (2022)
di: Levy, Mosh, et al.
Pubblicazione: (2022)
Evaluating the Robustness of the "Ensemble Everything Everywhere" Defense
di: Zhang, Jie, et al.
Pubblicazione: (2024)
di: Zhang, Jie, et al.
Pubblicazione: (2024)
Constructing Semantics-Aware Adversarial Examples with a Probabilistic Perspective
di: Zhang, Andi, et al.
Pubblicazione: (2023)
di: Zhang, Andi, et al.
Pubblicazione: (2023)
Differentially Private Adaptation of Diffusion Models via Noisy Aggregated Embeddings
di: Peetathawatchai, Pura, et al.
Pubblicazione: (2024)
di: Peetathawatchai, Pura, et al.
Pubblicazione: (2024)
Query-Based Adversarial Prompt Generation
di: Hayase, Jonathan, et al.
Pubblicazione: (2024)
di: Hayase, Jonathan, et al.
Pubblicazione: (2024)
Position: Towards Resilience Against Adversarial Examples
di: Dai, Sihui, et al.
Pubblicazione: (2024)
di: Dai, Sihui, et al.
Pubblicazione: (2024)
Universal Jailbreak Backdoors from Poisoned Human Feedback
di: Rando, Javier, et al.
Pubblicazione: (2023)
di: Rando, Javier, et al.
Pubblicazione: (2023)
Exploring and Mitigating Adversarial Manipulation of Voting-Based Leaderboards
di: Huang, Yangsibo, et al.
Pubblicazione: (2025)
di: Huang, Yangsibo, et al.
Pubblicazione: (2025)
Improved Generation of Adversarial Examples Against Safety-aligned LLMs
di: Li, Qizhang, et al.
Pubblicazione: (2024)
di: Li, Qizhang, et al.
Pubblicazione: (2024)
When and How to Fool Explainable Models (and Humans) with Adversarial Examples
di: Vadillo, Jon, et al.
Pubblicazione: (2021)
di: Vadillo, Jon, et al.
Pubblicazione: (2021)
Amatriciana: Exploiting Temporal GNNs for Robust and Efficient Money Laundering Detection
di: Di Gennaro, Marco, et al.
Pubblicazione: (2025)
di: Di Gennaro, Marco, et al.
Pubblicazione: (2025)
Certifiable Black-Box Attacks with Randomized Adversarial Examples: Breaking Defenses with Provable Confidence
di: Hong, Hanbin, et al.
Pubblicazione: (2023)
di: Hong, Hanbin, et al.
Pubblicazione: (2023)
Comprehensive Survey on Adversarial Examples in Cybersecurity: Impacts, Challenges, and Mitigation Strategies
di: Li, Li
Pubblicazione: (2024)
di: Li, Li
Pubblicazione: (2024)
Adversarial Perturbations Cannot Reliably Protect Artists From Generative AI
di: Hönig, Robert, et al.
Pubblicazione: (2024)
di: Hönig, Robert, et al.
Pubblicazione: (2024)
Intent Laundering: AI Safety Datasets Are Not What They Seem
di: Golchin, Shahriar, et al.
Pubblicazione: (2026)
di: Golchin, Shahriar, et al.
Pubblicazione: (2026)
Understanding Deep Learning defenses Against Adversarial Examples Through Visualizations for Dynamic Risk Assessment
di: Echeberria-Barrio, Xabier, et al.
Pubblicazione: (2024)
di: Echeberria-Barrio, Xabier, et al.
Pubblicazione: (2024)
AutoAdvExBench: Benchmarking autonomous exploitation of adversarial example defenses
di: Carlini, Nicholas, et al.
Pubblicazione: (2025)
di: Carlini, Nicholas, et al.
Pubblicazione: (2025)
Privacy-Preserving Graph-Based Machine Learning with Fully Homomorphic Encryption for Collaborative Anti-Money Laundering
di: Effendi, Fabrianne, et al.
Pubblicazione: (2024)
di: Effendi, Fabrianne, et al.
Pubblicazione: (2024)
Stealing the Invisible: Unveiling Pre-Trained CNN Models through Adversarial Examples and Timing Side-Channels
di: Shukla, Shubhi, et al.
Pubblicazione: (2024)
di: Shukla, Shubhi, et al.
Pubblicazione: (2024)
Constructing Adversarial Examples for Vertical Federated Learning: Optimal Client Corruption through Multi-Armed Bandit
di: Yao, Duanyi, et al.
Pubblicazione: (2024)
di: Yao, Duanyi, et al.
Pubblicazione: (2024)
DPxFin: Adaptive Differential Privacy for Anti-Money Laundering Detection via Reputation-Weighted Federated Learning
di: Kanagavelu, Renuga, et al.
Pubblicazione: (2026)
di: Kanagavelu, Renuga, et al.
Pubblicazione: (2026)
Privacy Side Channels in Machine Learning Systems
di: Debenedetti, Edoardo, et al.
Pubblicazione: (2023)
di: Debenedetti, Edoardo, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Adversarial ML Problems Are Getting Harder to Solve and to Evaluate
di: Rando, Javier, et al.
Pubblicazione: (2025) -
Rerouting LLM Routers
di: Shafran, Avital, et al.
Pubblicazione: (2025) -
Machine Against the RAG: Jamming Retrieval-Augmented Generation with Blocker Documents
di: Shafran, Avital, et al.
Pubblicazione: (2024) -
Adversarial Search Engine Optimization for Large Language Models
di: Nestaas, Fredrik, et al.
Pubblicazione: (2024) -
Beyond Labeling Oracles: What does it mean to steal ML models?
di: Shafran, Avital, et al.
Pubblicazione: (2023)