The Erasure Illusion: Stress-Testing the Generalization of LLM Forgetting Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jia, Hengrui, Li, Taoran, Guan, Jonas, Chandrasekaran, Varun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Challenges in Enabling Private Data Valuation
von: Fu, Yiwei, et al.
Veröffentlicht: (2026)
von: Fu, Yiwei, et al.
Veröffentlicht: (2026)
Layer-Targeted Multilingual Knowledge Erasure in Large Language Models
von: Li, Taoran, et al.
Veröffentlicht: (2026)
von: Li, Taoran, et al.
Veröffentlicht: (2026)
SoK: Understanding (New) Security Issues Across AI4Code Use Cases
von: Wu, Qilong, et al.
Veröffentlicht: (2025)
von: Wu, Qilong, et al.
Veröffentlicht: (2025)
Erased but Not Forgotten: How Backdoors Compromise Concept Erasure
von: Braun, Tobias, et al.
Veröffentlicht: (2025)
von: Braun, Tobias, et al.
Veröffentlicht: (2025)
Adversarial Illusions in Multi-Modal Embeddings
von: Zhang, Tingwei, et al.
Veröffentlicht: (2023)
von: Zhang, Tingwei, et al.
Veröffentlicht: (2023)
Gradients Look Alike: Sensitivity is Often Overestimated in DP-SGD
von: Thudi, Anvith, et al.
Veröffentlicht: (2023)
von: Thudi, Anvith, et al.
Veröffentlicht: (2023)
Rethinking the Vulnerability of Concept Erasure and a New Method
von: Richardson, Alex D., et al.
Veröffentlicht: (2025)
von: Richardson, Alex D., et al.
Veröffentlicht: (2025)
Backdoor Detection through Replicated Execution of Outsourced Training
von: Jia, Hengrui, et al.
Veröffentlicht: (2025)
von: Jia, Hengrui, et al.
Veröffentlicht: (2025)
AEGIS: Adversarial Target-Guided Retention-Data-Free Robust Concept Erasure from Diffusion Models
von: Li, Fengpeng, et al.
Veröffentlicht: (2026)
von: Li, Fengpeng, et al.
Veröffentlicht: (2026)
Learning to Forget using Hypernetworks
von: Rangel, Jose Miguel Lara, et al.
Veröffentlicht: (2024)
von: Rangel, Jose Miguel Lara, et al.
Veröffentlicht: (2024)
Fuzz-Testing Meets LLM-Based Agents: An Automated and Efficient Framework for Jailbreaking Text-To-Image Generation Models
von: Dong, Yingkai, et al.
Veröffentlicht: (2024)
von: Dong, Yingkai, et al.
Veröffentlicht: (2024)
AttackQA: Development and Adoption of a Dataset for Assisting Cybersecurity Operations using Fine-tuned and Open-Source LLMs
von: Krishna, Varun Badrinath
Veröffentlicht: (2024)
von: Krishna, Varun Badrinath
Veröffentlicht: (2024)
TracLLM: A Generic Framework for Attributing Long Context LLMs
von: Wang, Yanting, et al.
Veröffentlicht: (2025)
von: Wang, Yanting, et al.
Veröffentlicht: (2025)
ACU: Analytic Continual Unlearning for Efficient and Exact Forgetting with Privacy Preservation
von: Tang, Jianheng, et al.
Veröffentlicht: (2025)
von: Tang, Jianheng, et al.
Veröffentlicht: (2025)
An Empirical Evaluation of LLM-Generated Code Security Across Prompting Methods
von: Kharma, Mohammed, et al.
Veröffentlicht: (2026)
von: Kharma, Mohammed, et al.
Veröffentlicht: (2026)
Continual Learning with Strategic Selection and Forgetting for Network Intrusion Detection
von: Zhang, Xinchen, et al.
Veröffentlicht: (2024)
von: Zhang, Xinchen, et al.
Veröffentlicht: (2024)
Enhancing Continual Learning for Software Vulnerability Prediction: Addressing Catastrophic Forgetting via Hybrid-Confidence-Aware Selective Replay for Temporal LLM Fine-Tuning
von: Dou, Xuhui, et al.
Veröffentlicht: (2026)
von: Dou, Xuhui, et al.
Veröffentlicht: (2026)
Stateless Yet Not Forgetful: Implicit Memory as a Hidden Channel in LLMs
von: Salem, Ahmed, et al.
Veröffentlicht: (2026)
von: Salem, Ahmed, et al.
Veröffentlicht: (2026)
Data Unlearning Beyond Uniform Forgetting via Diffusion Time and Frequency Selection
von: Park, Jinseong, et al.
Veröffentlicht: (2025)
von: Park, Jinseong, et al.
Veröffentlicht: (2025)
Forget to Flourish: Leveraging Machine-Unlearning on Pretrained Language Models for Privacy Leakage
von: Rashid, Md Rafi Ur, et al.
Veröffentlicht: (2024)
von: Rashid, Md Rafi Ur, et al.
Veröffentlicht: (2024)
Odysseus: Jailbreaking Commercial Multimodal LLM-integrated Systems via Dual Steganography
von: Li, Songze, et al.
Veröffentlicht: (2025)
von: Li, Songze, et al.
Veröffentlicht: (2025)
BadSampler: Harnessing the Power of Catastrophic Forgetting to Poison Byzantine-robust Federated Learning
von: Liu, Yi, et al.
Veröffentlicht: (2024)
von: Liu, Yi, et al.
Veröffentlicht: (2024)
A Cognac Shot To Forget Bad Memories: Corrective Unlearning for Graph Neural Networks
von: Kolipaka, Varshita, et al.
Veröffentlicht: (2024)
von: Kolipaka, Varshita, et al.
Veröffentlicht: (2024)
Toward Principled LLM Safety Testing: Solving the Jailbreak Oracle Problem
von: Lin, Shuyi, et al.
Veröffentlicht: (2025)
von: Lin, Shuyi, et al.
Veröffentlicht: (2025)
Multi-Continental Healthcare Modelling Using Blockchain-Enabled Federated Learning
von: Sun, Rui, et al.
Veröffentlicht: (2024)
von: Sun, Rui, et al.
Veröffentlicht: (2024)
InfiCoEvalChain: A Blockchain-Based Decentralized Framework for Collaborative LLM Evaluation
von: Yang, Yifan, et al.
Veröffentlicht: (2026)
von: Yang, Yifan, et al.
Veröffentlicht: (2026)
Generating Synthetic Health Sensor Data for Privacy-Preserving Wearable Stress Detection
von: Lange, Lucas, et al.
Veröffentlicht: (2024)
von: Lange, Lucas, et al.
Veröffentlicht: (2024)
SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
von: Kutasov, Jonathan, et al.
Veröffentlicht: (2025)
von: Kutasov, Jonathan, et al.
Veröffentlicht: (2025)
TrafficLLM: Enhancing Large Language Models for Network Traffic Analysis with Generic Traffic Representation
von: Cui, Tianyu, et al.
Veröffentlicht: (2025)
von: Cui, Tianyu, et al.
Veröffentlicht: (2025)
Honeyval: A Comprehensive Evaluation Framework for LLM-powered HTTP Honeypots
von: Vero, Mark, et al.
Veröffentlicht: (2026)
von: Vero, Mark, et al.
Veröffentlicht: (2026)
Enhancing Reliability in LLM-Based Secure Code Generation
von: Kharma, Mohammed F., et al.
Veröffentlicht: (2026)
von: Kharma, Mohammed F., et al.
Veröffentlicht: (2026)
Bypassing LLM Watermarks with Color-Aware Substitutions
von: Wu, Qilong, et al.
Veröffentlicht: (2024)
von: Wu, Qilong, et al.
Veröffentlicht: (2024)
ACE: A Model Poisoning Attack on Contribution Evaluation Methods in Federated Learning
von: Xu, Zhangchen, et al.
Veröffentlicht: (2024)
von: Xu, Zhangchen, et al.
Veröffentlicht: (2024)
Privacy-Preserving Data Sharing in Agriculture: Enforcing Policy Rules for Secure and Confidential Data Synthesis
von: Kotal, Anantaa, et al.
Veröffentlicht: (2023)
von: Kotal, Anantaa, et al.
Veröffentlicht: (2023)
Systematic Scaling Analysis of Jailbreak Attacks in Large Language Models
von: Wang, Xiangwen, et al.
Veröffentlicht: (2026)
von: Wang, Xiangwen, et al.
Veröffentlicht: (2026)
Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMs
von: Panfilov, Alexander, et al.
Veröffentlicht: (2025)
von: Panfilov, Alexander, et al.
Veröffentlicht: (2025)
Evaluating Prompt Injection Defenses for Educational LLM Tutors: Security-Usability-Latency Trade-offs
von: Maiorano, Alexandre Cristovão
Veröffentlicht: (2026)
von: Maiorano, Alexandre Cristovão
Veröffentlicht: (2026)
Evaluating Differentially Private Generation of Domain-Specific Text
von: Sun, Yidan, et al.
Veröffentlicht: (2025)
von: Sun, Yidan, et al.
Veröffentlicht: (2025)
SynthCTI: LLM-Driven Synthetic CTI Generation to enhance MITRE Technique Mapping
von: Ruiz-Ródenas, Álvaro, et al.
Veröffentlicht: (2025)
von: Ruiz-Ródenas, Álvaro, et al.
Veröffentlicht: (2025)
The Autonomy Tax: Defense Training Breaks LLM Agents
von: Li, Shawn, et al.
Veröffentlicht: (2026)
von: Li, Shawn, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Challenges in Enabling Private Data Valuation
von: Fu, Yiwei, et al.
Veröffentlicht: (2026) -
Layer-Targeted Multilingual Knowledge Erasure in Large Language Models
von: Li, Taoran, et al.
Veröffentlicht: (2026) -
SoK: Understanding (New) Security Issues Across AI4Code Use Cases
von: Wu, Qilong, et al.
Veröffentlicht: (2025) -
Erased but Not Forgotten: How Backdoors Compromise Concept Erasure
von: Braun, Tobias, et al.
Veröffentlicht: (2025) -
Adversarial Illusions in Multi-Modal Embeddings
von: Zhang, Tingwei, et al.
Veröffentlicht: (2023)