LLM Ghostbusters: Surgical Hallucination Suppression via Adaptive Unlearning
Fuente:
arXiv
Guardado en:
| Autores principales: | Spracklen, Joseph, Aghazadeh, Pedram, Koushanfar, Farinaz, Jadliwala, Murtuza |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
We Have a Package for You! A Comprehensive Analysis of Package Hallucinations by Code Generating LLMs
por: Spracklen, Joseph, et al.
Publicado: (2024)
por: Spracklen, Joseph, et al.
Publicado: (2024)
Optimizing Privacy-Preserving Primitives to Support LLM-Scale Applications
por: Jandali, Yaman, et al.
Publicado: (2025)
por: Jandali, Yaman, et al.
Publicado: (2025)
Robust and Secure Code Watermarking for Large Language Models via ML/Crypto Codesign
por: Zhang, Ruisi, et al.
Publicado: (2025)
por: Zhang, Ruisi, et al.
Publicado: (2025)
Unlearned but Not Forgotten: Data Extraction after Exact Unlearning in LLM
por: Wu, Xiaoyu, et al.
Publicado: (2025)
por: Wu, Xiaoyu, et al.
Publicado: (2025)
Prompt and Circumstances: Evaluating the Efficacy of Human Prompt Inference in AI-Generated Art
por: Trinh, Khoi, et al.
Publicado: (2026)
por: Trinh, Khoi, et al.
Publicado: (2026)
From Domains to Instances: Dual-Granularity Data Synthesis for LLM Unlearning
por: Xu, Xiaoyu, et al.
Publicado: (2026)
por: Xu, Xiaoyu, et al.
Publicado: (2026)
REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations
por: Liang, Buyun, et al.
Publicado: (2026)
por: Liang, Buyun, et al.
Publicado: (2026)
SECA: Semantically Equivalent and Coherent Attacks for Eliciting LLM Hallucinations
por: Liang, Buyun, et al.
Publicado: (2025)
por: Liang, Buyun, et al.
Publicado: (2025)
VocalBridge: Latent Diffusion-Bridge Purification for Defeating Perturbation-Based Voiceprint Defenses
por: Abbasihafshejani, Maryam, et al.
Publicado: (2026)
por: Abbasihafshejani, Maryam, et al.
Publicado: (2026)
Textual Unlearning Gives a False Sense of Unlearning
por: Du, Jiacheng, et al.
Publicado: (2024)
por: Du, Jiacheng, et al.
Publicado: (2024)
Props for Machine-Learning Security
por: Juels, Ari, et al.
Publicado: (2024)
por: Juels, Ari, et al.
Publicado: (2024)
Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs
por: Xu, Xiaoyu, et al.
Publicado: (2025)
por: Xu, Xiaoyu, et al.
Publicado: (2025)
UnUnlearning: Unlearning is not sufficient for content regulation in advanced generative AI
por: Shumailov, Ilia, et al.
Publicado: (2024)
por: Shumailov, Ilia, et al.
Publicado: (2024)
EmMark: Robust Watermarks for IP Protection of Embedded Quantized Large Language Models
por: Zhang, Ruisi, et al.
Publicado: (2024)
por: Zhang, Ruisi, et al.
Publicado: (2024)
Watermarking Large Language Models and the Generated Content: Opportunities and Challenges
por: Zhang, Ruisi, et al.
Publicado: (2024)
por: Zhang, Ruisi, et al.
Publicado: (2024)
SAEs $\textit{Can}$ Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMs
por: Muhamed, Aashiq, et al.
Publicado: (2025)
por: Muhamed, Aashiq, et al.
Publicado: (2025)
Forgetting-MarI: LLM Unlearning via Marginal Information Regularization
por: Xu, Shizhou, et al.
Publicado: (2025)
por: Xu, Shizhou, et al.
Publicado: (2025)
Towards Robust Knowledge Unlearning: An Adversarial Framework for Assessing and Improving Unlearning Robustness in Large Language Models
por: Yuan, Hongbang, et al.
Publicado: (2024)
por: Yuan, Hongbang, et al.
Publicado: (2024)
Token-Specific Watermarking with Enhanced Detectability and Semantic Coherence for Large Language Models
por: Huo, Mingjia, et al.
Publicado: (2024)
por: Huo, Mingjia, et al.
Publicado: (2024)
Beyond Perplexity: A Lightweight Benchmark for Knowledge Retention in Supervised Fine-Tuning
por: Shabgahi, Soheil Zibakhsh, et al.
Publicado: (2026)
por: Shabgahi, Soheil Zibakhsh, et al.
Publicado: (2026)
An Adversarial Perspective on Machine Unlearning for AI Safety
por: Łucki, Jakub, et al.
Publicado: (2024)
por: Łucki, Jakub, et al.
Publicado: (2024)
Machine Unlearning of Pre-trained Large Language Models
por: Yao, Jin, et al.
Publicado: (2024)
por: Yao, Jin, et al.
Publicado: (2024)
Adaptive Instruction Composition for Automated LLM Red-Teaming
por: Zymet, Jesse, et al.
Publicado: (2026)
por: Zymet, Jesse, et al.
Publicado: (2026)
FIT to Forget: Robust Continual Unlearning for Large Language Models
por: Xu, Xiaoyu, et al.
Publicado: (2026)
por: Xu, Xiaoyu, et al.
Publicado: (2026)
OBLIVIATE: Robust and Practical Machine Unlearning for Large Language Models
por: Xu, Xiaoyu, et al.
Publicado: (2025)
por: Xu, Xiaoyu, et al.
Publicado: (2025)
Using Hallucinations to Bypass GPT4's Filter
por: Lemkin, Benjamin
Publicado: (2024)
por: Lemkin, Benjamin
Publicado: (2024)
Harry Potter is Still Here! Probing Knowledge Leakage in Targeted Unlearned Large Language Models via Automated Adversarial Prompting
por: To, Bang Trinh Tran, et al.
Publicado: (2025)
por: To, Bang Trinh Tran, et al.
Publicado: (2025)
Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM
por: Cao, Bochuan, et al.
Publicado: (2023)
por: Cao, Bochuan, et al.
Publicado: (2023)
In Vino Veritas and Vulnerabilities: Examining LLM Safety via Drunk Language Inducement
por: Shetty, Anudeex, et al.
Publicado: (2026)
por: Shetty, Anudeex, et al.
Publicado: (2026)
Prompt2Fingerprint: Plug-and-Play LLM Fingerprinting via Text-to-Weight Generation
por: Chen, Sixu, et al.
Publicado: (2026)
por: Chen, Sixu, et al.
Publicado: (2026)
REMARK-LLM: A Robust and Efficient Watermarking Framework for Generative Large Language Models
por: Zhang, Ruisi, et al.
Publicado: (2023)
por: Zhang, Ruisi, et al.
Publicado: (2023)
MergeGuard: Efficient Thwarting of Trojan Attacks in Machine Learning Models
por: Shabgahi, Soheil Zibakhsh, et al.
Publicado: (2025)
por: Shabgahi, Soheil Zibakhsh, et al.
Publicado: (2025)
Route to Rome Attack: Directing LLM Routers to Expensive Models via Adversarial Suffix Optimization
por: Tang, Haochun, et al.
Publicado: (2026)
por: Tang, Haochun, et al.
Publicado: (2026)
Agent-ToM: Learning to Monitor Autonomous LLM Agents via Theory-of-Mind Reasoning
por: Ahmed, Nesreen K., et al.
Publicado: (2026)
por: Ahmed, Nesreen K., et al.
Publicado: (2026)
Mark Your LLM: Detecting the Misuse of Open-Source Large Language Models via Watermarking
por: Xu, Yijie, et al.
Publicado: (2025)
por: Xu, Yijie, et al.
Publicado: (2025)
AttestLLM: Efficient Attestation Framework for Billion-scale On-device LLMs
por: Zhang, Ruisi, et al.
Publicado: (2025)
por: Zhang, Ruisi, et al.
Publicado: (2025)
IncogniText: Privacy-enhancing Conditional Text Anonymization via LLM-based Private Attribute Randomization
por: Frikha, Ahmed, et al.
Publicado: (2024)
por: Frikha, Ahmed, et al.
Publicado: (2024)
PII-Compass: Guiding LLM training data extraction prompts towards the target PII via grounding
por: Nakka, Krishna Kanth, et al.
Publicado: (2024)
por: Nakka, Krishna Kanth, et al.
Publicado: (2024)
SWaRL: Safeguard Code Watermarking via Reinforcement Learning
por: Javidnia, Neusha, et al.
Publicado: (2026)
por: Javidnia, Neusha, et al.
Publicado: (2026)
Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods
por: Jang, Yeonwoo, et al.
Publicado: (2025)
por: Jang, Yeonwoo, et al.
Publicado: (2025)
Ejemplares similares
-
We Have a Package for You! A Comprehensive Analysis of Package Hallucinations by Code Generating LLMs
por: Spracklen, Joseph, et al.
Publicado: (2024) -
Optimizing Privacy-Preserving Primitives to Support LLM-Scale Applications
por: Jandali, Yaman, et al.
Publicado: (2025) -
Robust and Secure Code Watermarking for Large Language Models via ML/Crypto Codesign
por: Zhang, Ruisi, et al.
Publicado: (2025) -
Unlearned but Not Forgotten: Data Extraction after Exact Unlearning in LLM
por: Wu, Xiaoyu, et al.
Publicado: (2025) -
Prompt and Circumstances: Evaluating the Efficacy of Human Prompt Inference in AI-Generated Art
por: Trinh, Khoi, et al.
Publicado: (2026)