Generalist++: A Meta-learning Framework for Mitigating Trade-off in Adversarial Training
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Yisen, Mo, Yichuan, Wang, Hongjun, Li, Junyi, Lin, Zhouchen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TERD: A Unified Framework for Safeguarding Diffusion Models Against Backdoors
by: Mo, Yichuan, et al.
Published: (2024)
by: Mo, Yichuan, et al.
Published: (2024)
Fight Back Against Jailbreaking via Prompt Adversarial Tuning
by: Mo, Yichuan, et al.
Published: (2024)
by: Mo, Yichuan, et al.
Published: (2024)
Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations
by: Wei, Zeming, et al.
Published: (2023)
by: Wei, Zeming, et al.
Published: (2023)
PID: Prompt-Independent Data Protection Against Latent Diffusion Models
by: Li, Ang, et al.
Published: (2024)
by: Li, Ang, et al.
Published: (2024)
Accuracy-Privacy Trade-off in the Mitigation of Membership Inference Attack in Federated Learning
by: Ahamed, Sayyed Farid, et al.
Published: (2024)
by: Ahamed, Sayyed Farid, et al.
Published: (2024)
MADE: Graph Backdoor Defense with Masked Unlearning
by: Lin, Xiao, et al.
Published: (2024)
by: Lin, Xiao, et al.
Published: (2024)
Identifying and Understanding Cross-Class Features in Adversarial Training
by: Wei, Zeming, et al.
Published: (2025)
by: Wei, Zeming, et al.
Published: (2025)
How to Craft Backdoors with Unlabeled Data Alone?
by: Wang, Yifei, et al.
Published: (2024)
by: Wang, Yifei, et al.
Published: (2024)
A Unified Learn-to-Distort-Data Framework for Privacy-Utility Trade-off in Trustworthy Federated Learning
by: Zhang, Xiaojin, et al.
Published: (2024)
by: Zhang, Xiaojin, et al.
Published: (2024)
Privacy-Utility Trade-off in Data Publication: A Bilevel Optimization Framework with Curvature-Guided Perturbation
by: Yin, Yi, et al.
Published: (2025)
by: Yin, Yi, et al.
Published: (2025)
Training on Fake Labels: Mitigating Label Leakage in Split Learning via Secure Dimension Transformation
by: Jiang, Yukun, et al.
Published: (2024)
by: Jiang, Yukun, et al.
Published: (2024)
On the Adversarial Transferability of Generalized "Skip Connections"
by: Wang, Yisen, et al.
Published: (2024)
by: Wang, Yisen, et al.
Published: (2024)
Mitigating the Structural Bias in Graph Adversarial Defenses
by: Fang, Junyuan, et al.
Published: (2025)
by: Fang, Junyuan, et al.
Published: (2025)
Clients Collaborate: Flexible Differentially Private Federated Learning with Guaranteed Improvement of Utility-Privacy Trade-off
by: Li, Yuecheng, et al.
Published: (2024)
by: Li, Yuecheng, et al.
Published: (2024)
Unveiling the Backdoor Mechanism Hidden Behind Catastrophic Overfitting in Fast Adversarial Training
by: Zhao, Mengnan, et al.
Published: (2026)
by: Zhao, Mengnan, et al.
Published: (2026)
GenoArmory: A Unified Evaluation Framework for Adversarial Attacks on Genomic Foundation Models
by: Luo, Haozheng, et al.
Published: (2025)
by: Luo, Haozheng, et al.
Published: (2025)
Improving Clean Accuracy via a Tangent-Space Perspective on Adversarial Training
by: Yi, Bongsoo, et al.
Published: (2024)
by: Yi, Bongsoo, et al.
Published: (2024)
MarkTune: Improving the Quality-Detectability Trade-off in Open-Weight LLM Watermarking
by: Zhao, Yizhou, et al.
Published: (2025)
by: Zhao, Yizhou, et al.
Published: (2025)
Mitigation of Camouflaged Adversarial Attacks in Autonomous Vehicles--A Case Study Using CARLA Simulator
by: Martinez, Yago Romano, et al.
Published: (2025)
by: Martinez, Yago Romano, et al.
Published: (2025)
Information Theoretic Adversarial Training of Large Language Models
by: Zhang, Yiwei, et al.
Published: (2026)
by: Zhang, Yiwei, et al.
Published: (2026)
EIA: Environmental Injection Attack on Generalist Web Agents for Privacy Leakage
by: Liao, Zeyi, et al.
Published: (2024)
by: Liao, Zeyi, et al.
Published: (2024)
Evaluating Prompt Injection Defenses for Educational LLM Tutors: Security-Usability-Latency Trade-offs
by: Maiorano, Alexandre Cristovão
Published: (2026)
by: Maiorano, Alexandre Cristovão
Published: (2026)
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval
by: Chen, Taiye, et al.
Published: (2025)
by: Chen, Taiye, et al.
Published: (2025)
A Novel Perturb-ability Score to Mitigate Evasion Adversarial Attacks on Flow-Based ML-NIDS
by: elShehaby, Mohamed, et al.
Published: (2024)
by: elShehaby, Mohamed, et al.
Published: (2024)
Disttack: Graph Adversarial Attacks Toward Distributed GNN Training
by: Zhang, Yuxiang, et al.
Published: (2024)
by: Zhang, Yuxiang, et al.
Published: (2024)
Embedding Hidden Adversarial Capabilities in Pre-Trained Diffusion Models
by: Beerens, Lucas, et al.
Published: (2025)
by: Beerens, Lucas, et al.
Published: (2025)
Defending Against Unforeseen Failure Modes with Latent Adversarial Training
by: Casper, Stephen, et al.
Published: (2024)
by: Casper, Stephen, et al.
Published: (2024)
Investigating the Impact of Quantization on Adversarial Robustness
by: Li, Qun, et al.
Published: (2024)
by: Li, Qun, et al.
Published: (2024)
NPAT Null-Space Projected Adversarial Training Towards Zero Deterioration
by: Hu, Hanyi, et al.
Published: (2024)
by: Hu, Hanyi, et al.
Published: (2024)
Adaptive PII Mitigation Framework for Large Language Models
by: Asthana, Shubhi, et al.
Published: (2025)
by: Asthana, Shubhi, et al.
Published: (2025)
Enhancing Security in Deep Reinforcement Learning: A Comprehensive Survey on Adversarial Attacks and Defenses
by: Yichao, Wu, et al.
Published: (2025)
by: Yichao, Wu, et al.
Published: (2025)
Machine-learned Adversarial Attacks against Fault Prediction Systems in Smart Electrical Grids
by: Ardito, Carmelo, et al.
Published: (2023)
by: Ardito, Carmelo, et al.
Published: (2023)
CTIGuardian: A Few-Shot Framework for Mitigating Privacy Leakage in Fine-Tuned LLMs
by: Arachchige, Shashie Dilhara Batan, et al.
Published: (2025)
by: Arachchige, Shashie Dilhara Batan, et al.
Published: (2025)
AEGIS: Adversarial Target-Guided Retention-Data-Free Robust Concept Erasure from Diffusion Models
by: Li, Fengpeng, et al.
Published: (2026)
by: Li, Fengpeng, et al.
Published: (2026)
Downstream Trade-offs of a Family of Text Watermarks
by: Ajith, Anirudh, et al.
Published: (2023)
by: Ajith, Anirudh, et al.
Published: (2023)
On the Duality Between Sharpness-Aware Minimization and Adversarial Training
by: Zhang, Yihao, et al.
Published: (2024)
by: Zhang, Yihao, et al.
Published: (2024)
Pandora's White-Box: Precise Training Data Detection and Extraction in Large Language Models
by: Wang, Jeffrey G., et al.
Published: (2024)
by: Wang, Jeffrey G., et al.
Published: (2024)
Untargeted Adversarial Attack on Knowledge Graph Embeddings
by: Zhao, Tianzhe, et al.
Published: (2024)
by: Zhao, Tianzhe, et al.
Published: (2024)
MoCo-EA: Exploiting Adversarial Mode Connectivity for Efficient Evolutionary Attacks
by: Kim, Hyo Seo, et al.
Published: (2026)
by: Kim, Hyo Seo, et al.
Published: (2026)
Variational Randomized Smoothing for Sample-Wise Adversarial Robustness
by: Hase, Ryo, et al.
Published: (2024)
by: Hase, Ryo, et al.
Published: (2024)
Similar Items
-
TERD: A Unified Framework for Safeguarding Diffusion Models Against Backdoors
by: Mo, Yichuan, et al.
Published: (2024) -
Fight Back Against Jailbreaking via Prompt Adversarial Tuning
by: Mo, Yichuan, et al.
Published: (2024) -
Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations
by: Wei, Zeming, et al.
Published: (2023) -
PID: Prompt-Independent Data Protection Against Latent Diffusion Models
by: Li, Ang, et al.
Published: (2024) -
Accuracy-Privacy Trade-off in the Mitigation of Membership Inference Attack in Federated Learning
by: Ahamed, Sayyed Farid, et al.
Published: (2024)