Jailbreaking the Non-Transferable Barrier via Test-Time Data Disguising
Fuente:
arXiv
Saved in:
| Main Authors: | Xiang, Yongli, Hong, Ziming, Yao, Lina, Wang, Dadong, Liu, Tongliang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Toward Robust Non-Transferable Learning: A Survey and Benchmark
by: Hong, Ziming, et al.
Published: (2025)
by: Hong, Ziming, et al.
Published: (2025)
When Data-Free Knowledge Distillation Meets Non-Transferable Teacher: Escaping Out-of-Distribution Trap is All You Need
by: Hong, Ziming, et al.
Published: (2025)
by: Hong, Ziming, et al.
Published: (2025)
AdLift: Lifting Adversarial Perturbations to Safeguard 3D Gaussian Splatting Assets Against Instruction-Driven Editing
by: Hong, Ziming, et al.
Published: (2025)
by: Hong, Ziming, et al.
Published: (2025)
MedBN: Robust Test-Time Adaptation against Malicious Test Samples
by: Park, Hyejin, et al.
Published: (2024)
by: Park, Hyejin, et al.
Published: (2024)
Intellectual Property Protection for 3D Gaussian Splatting Assets: A Survey
by: Zhao, Longjie, et al.
Published: (2026)
by: Zhao, Longjie, et al.
Published: (2026)
Interpretability-Guided Test-Time Adversarial Defense
by: Kulkarni, Akshay, et al.
Published: (2024)
by: Kulkarni, Akshay, et al.
Published: (2024)
Sonic: Fast and Transferable Data Poisoning on Clustering Algorithms
by: Villani, Francesco, et al.
Published: (2024)
by: Villani, Francesco, et al.
Published: (2024)
Boosting Adversarial Transferability via Residual Perturbation Attack
by: Peng, Jinjia, et al.
Published: (2025)
by: Peng, Jinjia, et al.
Published: (2025)
Improving Transferability of Adversarial Examples via Bayesian Attacks
by: Li, Qizhang, et al.
Published: (2023)
by: Li, Qizhang, et al.
Published: (2023)
Enabling Heterogeneous Adversarial Transferability via Feature Permutation Attacks
by: Wu, Tao, et al.
Published: (2025)
by: Wu, Tao, et al.
Published: (2025)
Towards Model Resistant to Transferable Adversarial Examples via Trigger Activation
by: Yu, Yi, et al.
Published: (2025)
by: Yu, Yi, et al.
Published: (2025)
JailbreakZoo: Survey, Landscapes, and Horizons in Jailbreaking Large Language and Vision-Language Models
by: Jin, Haibo, et al.
Published: (2024)
by: Jin, Haibo, et al.
Published: (2024)
Soften to Defend: Towards Adversarial Robustness via Self-Guided Label Refinement
by: Yu, Daiwei, et al.
Published: (2024)
by: Yu, Daiwei, et al.
Published: (2024)
Advancing Generalized Transfer Attack with Initialization Derived Bilevel Optimization and Dynamic Sequence Truncation
by: Liu, Yaohua, et al.
Published: (2024)
by: Liu, Yaohua, et al.
Published: (2024)
Multimodal Pragmatic Jailbreak on Text-to-image Models
by: Liu, Tong, et al.
Published: (2024)
by: Liu, Tong, et al.
Published: (2024)
R-TPT: Improving Adversarial Robustness of Vision-Language Models through Test-Time Prompt Tuning
by: Sheng, Lijun, et al.
Published: (2025)
by: Sheng, Lijun, et al.
Published: (2025)
Few-Shot Adversarial Prompt Learning on Vision-Language Models
by: Zhou, Yiwei, et al.
Published: (2024)
by: Zhou, Yiwei, et al.
Published: (2024)
Transferability Bound Theory: Exploring Relationship between Adversarial Transferability and Flatness
by: Fan, Mingyuan, et al.
Published: (2023)
by: Fan, Mingyuan, et al.
Published: (2023)
Jailbreaking Large Language Models Against Moderation Guardrails via Cipher Characters
by: Jin, Haibo, et al.
Published: (2024)
by: Jin, Haibo, et al.
Published: (2024)
Model Supply Chain Poisoning: Backdooring Pre-trained Models via Embedding Indistinguishability
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
FC-Attack: Jailbreaking Multimodal Large Language Models via Auto-Generated Flowcharts
by: Zhang, Ziyi, et al.
Published: (2025)
by: Zhang, Ziyi, et al.
Published: (2025)
Improving Integrated Gradient-based Transferable Adversarial Examples by Refining the Integration Path
by: Ren, Yuchen, et al.
Published: (2024)
by: Ren, Yuchen, et al.
Published: (2024)
On the Privacy Effect of Data Enhancement via the Lens of Memorization
by: Li, Xiao, et al.
Published: (2022)
by: Li, Xiao, et al.
Published: (2022)
Deep Learning with Data Privacy via Residual Perturbation
by: Tao, Wenqi, et al.
Published: (2024)
by: Tao, Wenqi, et al.
Published: (2024)
Transferable Adversarial Examples with Bayes Approach
by: Fan, Mingyuan, et al.
Published: (2022)
by: Fan, Mingyuan, et al.
Published: (2022)
Failures to Find Transferable Image Jailbreaks Between Vision-Language Models
by: Schaeffer, Rylan, et al.
Published: (2024)
by: Schaeffer, Rylan, et al.
Published: (2024)
Sparse and Transferable Universal Singular Vectors Attack
by: Kuvshinova, Kseniia, et al.
Published: (2024)
by: Kuvshinova, Kseniia, et al.
Published: (2024)
Why Does Little Robustness Help? A Further Step Towards Understanding Adversarial Transferability
by: Zhang, Yechao, et al.
Published: (2023)
by: Zhang, Yechao, et al.
Published: (2023)
Model Inversion Robustness: Can Transfer Learning Help?
by: Ho, Sy-Tuyen, et al.
Published: (2024)
by: Ho, Sy-Tuyen, et al.
Published: (2024)
Transferable Adversarial Attacks on SAM and Its Downstream Models
by: Xia, Song, et al.
Published: (2024)
by: Xia, Song, et al.
Published: (2024)
Jailbreaking Attack against Multimodal Large Language Model
by: Niu, Zhenxing, et al.
Published: (2024)
by: Niu, Zhenxing, et al.
Published: (2024)
Differentially Private Synthetic Data via Foundation Model APIs 1: Images
by: Lin, Zinan, et al.
Published: (2023)
by: Lin, Zinan, et al.
Published: (2023)
R-CONV: An Analytical Approach for Efficient Data Reconstruction via Convolutional Gradients
by: Eltaras, Tamer Ahmed, et al.
Published: (2024)
by: Eltaras, Tamer Ahmed, et al.
Published: (2024)
Revisiting Transferable Adversarial Images: Systemization, Evaluation, and New Insights
by: Zhao, Zhengyu, et al.
Published: (2023)
by: Zhao, Zhengyu, et al.
Published: (2023)
Privacy-Preserving CNN Training with Transfer Learning: Multiclass Logistic Regression
by: Chiang, John
Published: (2023)
by: Chiang, John
Published: (2023)
Risks When Sharing LoRA Fine-Tuned Diffusion Model Weights
by: Yao, Dixi
Published: (2024)
by: Yao, Dixi
Published: (2024)
Differentially Private Synthetic Data via APIs 3: Using Simulators Instead of Foundation Model
by: Lin, Zinan, et al.
Published: (2025)
by: Lin, Zinan, et al.
Published: (2025)
Exploring Adversarial Attacks against Latent Diffusion Model from the Perspective of Adversarial Transferability
by: Chen, Junxi, et al.
Published: (2024)
by: Chen, Junxi, et al.
Published: (2024)
VERA-V: Variational Inference Framework for Jailbreaking Vision-Language Models
by: Liao, Qilin, et al.
Published: (2025)
by: Liao, Qilin, et al.
Published: (2025)
GASP: Efficient Black-Box Generation of Adversarial Suffixes for Jailbreaking LLMs
by: Basani, Advik Raj, et al.
Published: (2024)
by: Basani, Advik Raj, et al.
Published: (2024)
Similar Items
-
Toward Robust Non-Transferable Learning: A Survey and Benchmark
by: Hong, Ziming, et al.
Published: (2025) -
When Data-Free Knowledge Distillation Meets Non-Transferable Teacher: Escaping Out-of-Distribution Trap is All You Need
by: Hong, Ziming, et al.
Published: (2025) -
AdLift: Lifting Adversarial Perturbations to Safeguard 3D Gaussian Splatting Assets Against Instruction-Driven Editing
by: Hong, Ziming, et al.
Published: (2025) -
MedBN: Robust Test-Time Adaptation against Malicious Test Samples
by: Park, Hyejin, et al.
Published: (2024) -
Intellectual Property Protection for 3D Gaussian Splatting Assets: A Survey
by: Zhao, Longjie, et al.
Published: (2026)