The Boy Who Survived: Removing Harry Potter from an LLM is harder than reported
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Shostack, Adam |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Harry Potter is Still Here! Probing Knowledge Leakage in Targeted Unlearned Large Language Models via Automated Adversarial Prompting
von: To, Bang Trinh Tran, et al.
Veröffentlicht: (2025)
von: To, Bang Trinh Tran, et al.
Veröffentlicht: (2025)
Who can we trust? LLM-as-a-jury for Comparative Assessment
von: Qian, Mengjie, et al.
Veröffentlicht: (2026)
von: Qian, Mengjie, et al.
Veröffentlicht: (2026)
The illusion of a perfect metric: Why evaluating AI's words is harder than it looks
von: Oliva, Maria Paz, et al.
Veröffentlicht: (2025)
von: Oliva, Maria Paz, et al.
Veröffentlicht: (2025)
Who Gets Which Message? Auditing Demographic Bias in LLM-Generated Targeted Text
von: Islam, Tunazzina
Veröffentlicht: (2026)
von: Islam, Tunazzina
Veröffentlicht: (2026)
Robustly Improving LLM Fairness in Realistic Settings via Interpretability
von: Karvonen, Adam, et al.
Veröffentlicht: (2025)
von: Karvonen, Adam, et al.
Veröffentlicht: (2025)
Does Machine Unlearning Truly Remove Knowledge?
von: Chen, Haokun, et al.
Veröffentlicht: (2025)
von: Chen, Haokun, et al.
Veröffentlicht: (2025)
Mechanism of Task-oriented Information Removal in In-context Learning
von: Cho, Hakaze, et al.
Veröffentlicht: (2025)
von: Cho, Hakaze, et al.
Veröffentlicht: (2025)
Language models are better than humans at next-token prediction
von: Shlegeris, Buck, et al.
Veröffentlicht: (2022)
von: Shlegeris, Buck, et al.
Veröffentlicht: (2022)
R.I.P.: Better Models by Survival of the Fittest Prompts
von: Yu, Ping, et al.
Veröffentlicht: (2025)
von: Yu, Ping, et al.
Veröffentlicht: (2025)
AdapterSwap: Continuous Training of LLMs with Data Removal and Access-Control Guarantees
von: Fleshman, William, et al.
Veröffentlicht: (2024)
von: Fleshman, William, et al.
Veröffentlicht: (2024)
Repeat After Me: Transformers are Better than State Space Models at Copying
von: Jelassi, Samy, et al.
Veröffentlicht: (2024)
von: Jelassi, Samy, et al.
Veröffentlicht: (2024)
BiasGym: A Simple and Generalizable Framework for Analyzing and Removing Biases through Elicitation
von: Islam, Sekh Mainul, et al.
Veröffentlicht: (2025)
von: Islam, Sekh Mainul, et al.
Veröffentlicht: (2025)
BED-LLM: Intelligent Information Gathering with LLMs and Bayesian Experimental Design
von: Choudhury, Deepro, et al.
Veröffentlicht: (2025)
von: Choudhury, Deepro, et al.
Veröffentlicht: (2025)
Two are better than one: Context window extension with multi-grained self-injection
von: Han, Wei, et al.
Veröffentlicht: (2024)
von: Han, Wei, et al.
Veröffentlicht: (2024)
Revisiting Chain-of-Thought Prompting: Zero-shot Can Be Stronger than Few-shot
von: Cheng, Xiang, et al.
Veröffentlicht: (2025)
von: Cheng, Xiang, et al.
Veröffentlicht: (2025)
Mechanistic origins of catastrophic forgetting: why RL preserves circuits better than SFT?
von: Nunez, Jeanmely Rojas, et al.
Veröffentlicht: (2026)
von: Nunez, Jeanmely Rojas, et al.
Veröffentlicht: (2026)
Beyond the Last Answer: Your Reasoning Trace Uncovers More than You Think
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2025)
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2025)
Iterative Label Refinement Matters More than Preference Optimization under Weak Supervision
von: Ye, Yaowen, et al.
Veröffentlicht: (2025)
von: Ye, Yaowen, et al.
Veröffentlicht: (2025)
Time-To-Inconsistency: A Survival Analysis of Large Language Model Robustness to Adversarial Attacks
von: Li, Yubo, et al.
Veröffentlicht: (2025)
von: Li, Yubo, et al.
Veröffentlicht: (2025)
Mitigating LLM Hallucinations via Conformal Abstention
von: Yadkori, Yasin Abbasi, et al.
Veröffentlicht: (2024)
von: Yadkori, Yasin Abbasi, et al.
Veröffentlicht: (2024)
MoEs Are Stronger than You Think: Hyper-Parallel Inference Scaling with RoE
von: Zibakhsh, Soheil, et al.
Veröffentlicht: (2025)
von: Zibakhsh, Soheil, et al.
Veröffentlicht: (2025)
Bridging the Gap: Enhancing LLM Performance for Low-Resource African Languages with New Benchmarks, Fine-Tuning, and Cultural Adjustments
von: Alhanai, Tuka, et al.
Veröffentlicht: (2024)
von: Alhanai, Tuka, et al.
Veröffentlicht: (2024)
OptiMer: Optimal Distribution Vector Merging Is Better than Data Mixing for Continual Pre-Training
von: Song, Haiyue, et al.
Veröffentlicht: (2026)
von: Song, Haiyue, et al.
Veröffentlicht: (2026)
Who Benefits From Sinus Surgery? Comparing Generative AI and Supervised Machine Learning for Predicting Surgical Outcomes in Chronic Rhinosinusitis
von: Chowdhury, Sayeed Shafayet, et al.
Veröffentlicht: (2026)
von: Chowdhury, Sayeed Shafayet, et al.
Veröffentlicht: (2026)
LLM Cyber Evaluations Don't Capture Real-World Risk
von: Lukošiūtė, Kamilė, et al.
Veröffentlicht: (2025)
von: Lukošiūtė, Kamilė, et al.
Veröffentlicht: (2025)
IITK at SemEval-2024 Task 10: Who is the speaker? Improving Emotion Recognition and Flip Reasoning in Conversations via Speaker Embeddings
von: Patel, Shubham, et al.
Veröffentlicht: (2024)
von: Patel, Shubham, et al.
Veröffentlicht: (2024)
Set-LLM: A Permutation-Invariant LLM
von: Egressy, Beni, et al.
Veröffentlicht: (2025)
von: Egressy, Beni, et al.
Veröffentlicht: (2025)
LLM Chemistry Estimation for Multi-LLM Recommendation
von: Sanchez, Huascar, et al.
Veröffentlicht: (2025)
von: Sanchez, Huascar, et al.
Veröffentlicht: (2025)
Removing Spurious Correlation from Neural Network Interpretations
von: Fotouhi, Milad, et al.
Veröffentlicht: (2024)
von: Fotouhi, Milad, et al.
Veröffentlicht: (2024)
User-LLM: Efficient LLM Contextualization with User Embeddings
von: Ning, Lin, et al.
Veröffentlicht: (2024)
von: Ning, Lin, et al.
Veröffentlicht: (2024)
Reactive Transformer (RxT) -- Stateful Real-Time Processing for Event-Driven Reactive Language Models
von: Filipek, Adam
Veröffentlicht: (2025)
von: Filipek, Adam
Veröffentlicht: (2025)
LogProber: Disentangling confidence from contamination in LLM responses
von: Yax, Nicolas, et al.
Veröffentlicht: (2024)
von: Yax, Nicolas, et al.
Veröffentlicht: (2024)
VBART: The Turkish LLM
von: Turker, Meliksah, et al.
Veröffentlicht: (2024)
von: Turker, Meliksah, et al.
Veröffentlicht: (2024)
LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
von: Fu, Qichen, et al.
Veröffentlicht: (2024)
von: Fu, Qichen, et al.
Veröffentlicht: (2024)
BPO: Staying Close to the Behavior LLM Creates Better Online LLM Alignment
von: Xu, Wenda, et al.
Veröffentlicht: (2024)
von: Xu, Wenda, et al.
Veröffentlicht: (2024)
Large Language Models Assume People are More Rational than We Really are
von: Liu, Ryan, et al.
Veröffentlicht: (2024)
von: Liu, Ryan, et al.
Veröffentlicht: (2024)
LLM Maybe LongLM: Self-Extend LLM Context Window Without Tuning
von: Jin, Hongye, et al.
Veröffentlicht: (2024)
von: Jin, Hongye, et al.
Veröffentlicht: (2024)
LLM See, LLM Do: Guiding Data Generation to Target Non-Differentiable Objectives
von: Shimabucoro, Luísa, et al.
Veröffentlicht: (2024)
von: Shimabucoro, Luísa, et al.
Veröffentlicht: (2024)
MixEval: Deriving Wisdom of the Crowd from LLM Benchmark Mixtures
von: Ni, Jinjie, et al.
Veröffentlicht: (2024)
von: Ni, Jinjie, et al.
Veröffentlicht: (2024)
R-Zero: Self-Evolving Reasoning LLM from Zero Data
von: Huang, Chengsong, et al.
Veröffentlicht: (2025)
von: Huang, Chengsong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Harry Potter is Still Here! Probing Knowledge Leakage in Targeted Unlearned Large Language Models via Automated Adversarial Prompting
von: To, Bang Trinh Tran, et al.
Veröffentlicht: (2025) -
Who can we trust? LLM-as-a-jury for Comparative Assessment
von: Qian, Mengjie, et al.
Veröffentlicht: (2026) -
The illusion of a perfect metric: Why evaluating AI's words is harder than it looks
von: Oliva, Maria Paz, et al.
Veröffentlicht: (2025) -
Who Gets Which Message? Auditing Demographic Bias in LLM-Generated Targeted Text
von: Islam, Tunazzina
Veröffentlicht: (2026) -
Robustly Improving LLM Fairness in Realistic Settings via Interpretability
von: Karvonen, Adam, et al.
Veröffentlicht: (2025)