Gespeichert in:
| Hauptverfasser: | Rezkellah, Fatmazohra, Dakhmouche, Ramzi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2510.03567 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Adversarial Representation Engineering: A General Model Editing Framework for Large Language Models
von: Zhang, Yihao, et al.
Veröffentlicht: (2024)
von: Zhang, Yihao, et al.
Veröffentlicht: (2024)
Exploring the Robustness of In-Context Learning with Noisy Labels
von: Cheng, Chen, et al.
Veröffentlicht: (2024)
von: Cheng, Chen, et al.
Veröffentlicht: (2024)
The Utility and Complexity of in- and out-of-Distribution Machine Unlearning
von: Allouah, Youssef, et al.
Veröffentlicht: (2024)
von: Allouah, Youssef, et al.
Veröffentlicht: (2024)
Secure LLM Fine-Tuning via Safety-Aware Probing
von: Wu, Chengcan, et al.
Veröffentlicht: (2025)
von: Wu, Chengcan, et al.
Veröffentlicht: (2025)
Dynamic Orthogonal Continual Fine-tuning for Mitigating Catastrophic Forgettings
von: Zhang, Zhixin, et al.
Veröffentlicht: (2025)
von: Zhang, Zhixin, et al.
Veröffentlicht: (2025)
Boosting Jailbreak Attack with Momentum
von: Zhang, Yihao, et al.
Veröffentlicht: (2024)
von: Zhang, Yihao, et al.
Veröffentlicht: (2024)
RAPO: Risk-Aware Preference Optimization for Generalizable Safe Reasoning
von: Wei, Zeming, et al.
Veröffentlicht: (2026)
von: Wei, Zeming, et al.
Veröffentlicht: (2026)
GSE: Group-wise Sparse and Explainable Adversarial Attacks
von: Sadiku, Shpresim, et al.
Veröffentlicht: (2023)
von: Sadiku, Shpresim, et al.
Veröffentlicht: (2023)
UCD: Unlearning in LLMs via Contrastive Decoding
von: Suriyakumar, Vinith M., et al.
Veröffentlicht: (2025)
von: Suriyakumar, Vinith M., et al.
Veröffentlicht: (2025)
An Adversarial Perspective on Machine Unlearning for AI Safety
von: Łucki, Jakub, et al.
Veröffentlicht: (2024)
von: Łucki, Jakub, et al.
Veröffentlicht: (2024)
Differential Privacy via Distributionally Robust Optimization
von: Selvi, Aras, et al.
Veröffentlicht: (2023)
von: Selvi, Aras, et al.
Veröffentlicht: (2023)
Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs
von: Xu, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Xu, Xiaoyu, et al.
Veröffentlicht: (2025)
SMI: Statistical Membership Inference for Reliable Unlearned Model Auditing
von: Sun, Jialong, et al.
Veröffentlicht: (2026)
von: Sun, Jialong, et al.
Veröffentlicht: (2026)
Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods
von: Jang, Yeonwoo, et al.
Veröffentlicht: (2025)
von: Jang, Yeonwoo, et al.
Veröffentlicht: (2025)
Towards Robust Knowledge Unlearning: An Adversarial Framework for Assessing and Improving Unlearning Robustness in Large Language Models
von: Yuan, Hongbang, et al.
Veröffentlicht: (2024)
von: Yuan, Hongbang, et al.
Veröffentlicht: (2024)
Machine Unlearning: Taxonomy, Metrics, Applications, Challenges, and Prospects
von: Li, Na, et al.
Veröffentlicht: (2024)
von: Li, Na, et al.
Veröffentlicht: (2024)
Efficient Optimization Algorithms for Linear Adversarial Training
von: RIbeiro, Antônio H., et al.
Veröffentlicht: (2024)
von: RIbeiro, Antônio H., et al.
Veröffentlicht: (2024)
Expected Harm: Rethinking Safety Evaluation of (Mis)Aligned LLMs
von: Chen, Yen-Shan, et al.
Veröffentlicht: (2026)
von: Chen, Yen-Shan, et al.
Veröffentlicht: (2026)
Second-Order Min-Max Optimization with Lazy Hessians
von: Chen, Lesi, et al.
Veröffentlicht: (2024)
von: Chen, Lesi, et al.
Veröffentlicht: (2024)
Kernel Learning with Adversarial Features: Numerical Efficiency and Adaptive Regularization
von: Ribeiro, Antônio H., et al.
Veröffentlicht: (2025)
von: Ribeiro, Antônio H., et al.
Veröffentlicht: (2025)
LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet
von: Li, Nathaniel, et al.
Veröffentlicht: (2024)
von: Li, Nathaniel, et al.
Veröffentlicht: (2024)
k-SemStamp: A Clustering-Based Semantic Watermark for Detection of Machine-Generated Text
von: Hou, Abe Bohan, et al.
Veröffentlicht: (2024)
von: Hou, Abe Bohan, et al.
Veröffentlicht: (2024)
SimMark: A Robust Sentence-Level Similarity-Based Watermarking Algorithm for Large Language Models
von: Dabiriaghdam, Amirhossein, et al.
Veröffentlicht: (2025)
von: Dabiriaghdam, Amirhossein, et al.
Veröffentlicht: (2025)
Identifying and Understanding Cross-Class Features in Adversarial Training
von: Wei, Zeming, et al.
Veröffentlicht: (2025)
von: Wei, Zeming, et al.
Veröffentlicht: (2025)
Byzantine-Robust and Differentially Private Federated Optimization under Weaker Assumptions
von: Islamov, Rustem, et al.
Veröffentlicht: (2026)
von: Islamov, Rustem, et al.
Veröffentlicht: (2026)
OBLIVIATE: Robust and Practical Machine Unlearning for Large Language Models
von: Xu, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Xu, Xiaoyu, et al.
Veröffentlicht: (2025)
Calibrated Adversarial Sampling: Multi-Armed Bandit-Guided Generalization Against Unforeseen Attacks
von: Wang, Rui, et al.
Veröffentlicht: (2025)
von: Wang, Rui, et al.
Veröffentlicht: (2025)
Machine Unlearning Fails to Remove Data Poisoning Attacks
von: Pawelczyk, Martin, et al.
Veröffentlicht: (2024)
von: Pawelczyk, Martin, et al.
Veröffentlicht: (2024)
Hidden Poison: Machine Unlearning Enables Camouflaged Poisoning Attacks
von: Di, Jimmy Z., et al.
Veröffentlicht: (2022)
von: Di, Jimmy Z., et al.
Veröffentlicht: (2022)
On the Duality Between Sharpness-Aware Minimization and Adversarial Training
von: Zhang, Yihao, et al.
Veröffentlicht: (2024)
von: Zhang, Yihao, et al.
Veröffentlicht: (2024)
Smoothed Normalization for Efficient Distributed Private Optimization
von: Shulgin, Egor, et al.
Veröffentlicht: (2025)
von: Shulgin, Egor, et al.
Veröffentlicht: (2025)
Residual-Evasive Attacks on ADMM in Distributed Optimization
von: Bruckmeier, Sabrina, et al.
Veröffentlicht: (2025)
von: Bruckmeier, Sabrina, et al.
Veröffentlicht: (2025)
FedADMM-InSa: An Inexact and Self-Adaptive ADMM for Federated Learning
von: Song, Yongcun, et al.
Veröffentlicht: (2024)
von: Song, Yongcun, et al.
Veröffentlicht: (2024)
The Privacy Power of Correlated Noise in Decentralized Learning
von: Allouah, Youssef, et al.
Veröffentlicht: (2024)
von: Allouah, Youssef, et al.
Veröffentlicht: (2024)
SafeCOMM: A Study on Safety Degradation in Fine-Tuned Telecom Large Language Models
von: Djuhera, Aladin, et al.
Veröffentlicht: (2025)
von: Djuhera, Aladin, et al.
Veröffentlicht: (2025)
BadFair: Backdoored Fairness Attacks with Group-conditioned Triggers
von: Xue, Jiaqi, et al.
Veröffentlicht: (2024)
von: Xue, Jiaqi, et al.
Veröffentlicht: (2024)
Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models
von: An, Bang, et al.
Veröffentlicht: (2024)
von: An, Bang, et al.
Veröffentlicht: (2024)
In the Name of Fairness: Assessing the Bias in Clinical Record De-identification
von: Xiao, Yuxin, et al.
Veröffentlicht: (2023)
von: Xiao, Yuxin, et al.
Veröffentlicht: (2023)
PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts
von: Zhu, Kaijie, et al.
Veröffentlicht: (2023)
von: Zhu, Kaijie, et al.
Veröffentlicht: (2023)
Self-Evaluation as a Defense Against Adversarial Attacks on LLMs
von: Brown, Hannah, et al.
Veröffentlicht: (2024)
von: Brown, Hannah, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Adversarial Representation Engineering: A General Model Editing Framework for Large Language Models
von: Zhang, Yihao, et al.
Veröffentlicht: (2024) -
Exploring the Robustness of In-Context Learning with Noisy Labels
von: Cheng, Chen, et al.
Veröffentlicht: (2024) -
The Utility and Complexity of in- and out-of-Distribution Machine Unlearning
von: Allouah, Youssef, et al.
Veröffentlicht: (2024) -
Secure LLM Fine-Tuning via Safety-Aware Probing
von: Wu, Chengcan, et al.
Veröffentlicht: (2025) -
Dynamic Orthogonal Continual Fine-tuning for Mitigating Catastrophic Forgettings
von: Zhang, Zhixin, et al.
Veröffentlicht: (2025)