The Boy Who Survived: Removing Harry Potter from an LLM is harder than reported
Fuente:
arXiv
Salvato in:
| Autore principale: | Shostack, Adam |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Harry Potter is Still Here! Probing Knowledge Leakage in Targeted Unlearned Large Language Models via Automated Adversarial Prompting
di: To, Bang Trinh Tran, et al.
Pubblicazione: (2025)
di: To, Bang Trinh Tran, et al.
Pubblicazione: (2025)
Who can we trust? LLM-as-a-jury for Comparative Assessment
di: Qian, Mengjie, et al.
Pubblicazione: (2026)
di: Qian, Mengjie, et al.
Pubblicazione: (2026)
The illusion of a perfect metric: Why evaluating AI's words is harder than it looks
di: Oliva, Maria Paz, et al.
Pubblicazione: (2025)
di: Oliva, Maria Paz, et al.
Pubblicazione: (2025)
Who Gets Which Message? Auditing Demographic Bias in LLM-Generated Targeted Text
di: Islam, Tunazzina
Pubblicazione: (2026)
di: Islam, Tunazzina
Pubblicazione: (2026)
Robustly Improving LLM Fairness in Realistic Settings via Interpretability
di: Karvonen, Adam, et al.
Pubblicazione: (2025)
di: Karvonen, Adam, et al.
Pubblicazione: (2025)
Does Machine Unlearning Truly Remove Knowledge?
di: Chen, Haokun, et al.
Pubblicazione: (2025)
di: Chen, Haokun, et al.
Pubblicazione: (2025)
Mechanism of Task-oriented Information Removal in In-context Learning
di: Cho, Hakaze, et al.
Pubblicazione: (2025)
di: Cho, Hakaze, et al.
Pubblicazione: (2025)
Language models are better than humans at next-token prediction
di: Shlegeris, Buck, et al.
Pubblicazione: (2022)
di: Shlegeris, Buck, et al.
Pubblicazione: (2022)
R.I.P.: Better Models by Survival of the Fittest Prompts
di: Yu, Ping, et al.
Pubblicazione: (2025)
di: Yu, Ping, et al.
Pubblicazione: (2025)
AdapterSwap: Continuous Training of LLMs with Data Removal and Access-Control Guarantees
di: Fleshman, William, et al.
Pubblicazione: (2024)
di: Fleshman, William, et al.
Pubblicazione: (2024)
Repeat After Me: Transformers are Better than State Space Models at Copying
di: Jelassi, Samy, et al.
Pubblicazione: (2024)
di: Jelassi, Samy, et al.
Pubblicazione: (2024)
BiasGym: A Simple and Generalizable Framework for Analyzing and Removing Biases through Elicitation
di: Islam, Sekh Mainul, et al.
Pubblicazione: (2025)
di: Islam, Sekh Mainul, et al.
Pubblicazione: (2025)
BED-LLM: Intelligent Information Gathering with LLMs and Bayesian Experimental Design
di: Choudhury, Deepro, et al.
Pubblicazione: (2025)
di: Choudhury, Deepro, et al.
Pubblicazione: (2025)
Two are better than one: Context window extension with multi-grained self-injection
di: Han, Wei, et al.
Pubblicazione: (2024)
di: Han, Wei, et al.
Pubblicazione: (2024)
Revisiting Chain-of-Thought Prompting: Zero-shot Can Be Stronger than Few-shot
di: Cheng, Xiang, et al.
Pubblicazione: (2025)
di: Cheng, Xiang, et al.
Pubblicazione: (2025)
Mechanistic origins of catastrophic forgetting: why RL preserves circuits better than SFT?
di: Nunez, Jeanmely Rojas, et al.
Pubblicazione: (2026)
di: Nunez, Jeanmely Rojas, et al.
Pubblicazione: (2026)
Beyond the Last Answer: Your Reasoning Trace Uncovers More than You Think
di: Hammoud, Hasan Abed Al Kader, et al.
Pubblicazione: (2025)
di: Hammoud, Hasan Abed Al Kader, et al.
Pubblicazione: (2025)
Iterative Label Refinement Matters More than Preference Optimization under Weak Supervision
di: Ye, Yaowen, et al.
Pubblicazione: (2025)
di: Ye, Yaowen, et al.
Pubblicazione: (2025)
Time-To-Inconsistency: A Survival Analysis of Large Language Model Robustness to Adversarial Attacks
di: Li, Yubo, et al.
Pubblicazione: (2025)
di: Li, Yubo, et al.
Pubblicazione: (2025)
Mitigating LLM Hallucinations via Conformal Abstention
di: Yadkori, Yasin Abbasi, et al.
Pubblicazione: (2024)
di: Yadkori, Yasin Abbasi, et al.
Pubblicazione: (2024)
MoEs Are Stronger than You Think: Hyper-Parallel Inference Scaling with RoE
di: Zibakhsh, Soheil, et al.
Pubblicazione: (2025)
di: Zibakhsh, Soheil, et al.
Pubblicazione: (2025)
Bridging the Gap: Enhancing LLM Performance for Low-Resource African Languages with New Benchmarks, Fine-Tuning, and Cultural Adjustments
di: Alhanai, Tuka, et al.
Pubblicazione: (2024)
di: Alhanai, Tuka, et al.
Pubblicazione: (2024)
OptiMer: Optimal Distribution Vector Merging Is Better than Data Mixing for Continual Pre-Training
di: Song, Haiyue, et al.
Pubblicazione: (2026)
di: Song, Haiyue, et al.
Pubblicazione: (2026)
Who Benefits From Sinus Surgery? Comparing Generative AI and Supervised Machine Learning for Predicting Surgical Outcomes in Chronic Rhinosinusitis
di: Chowdhury, Sayeed Shafayet, et al.
Pubblicazione: (2026)
di: Chowdhury, Sayeed Shafayet, et al.
Pubblicazione: (2026)
LLM Cyber Evaluations Don't Capture Real-World Risk
di: Lukošiūtė, Kamilė, et al.
Pubblicazione: (2025)
di: Lukošiūtė, Kamilė, et al.
Pubblicazione: (2025)
IITK at SemEval-2024 Task 10: Who is the speaker? Improving Emotion Recognition and Flip Reasoning in Conversations via Speaker Embeddings
di: Patel, Shubham, et al.
Pubblicazione: (2024)
di: Patel, Shubham, et al.
Pubblicazione: (2024)
Set-LLM: A Permutation-Invariant LLM
di: Egressy, Beni, et al.
Pubblicazione: (2025)
di: Egressy, Beni, et al.
Pubblicazione: (2025)
LLM Chemistry Estimation for Multi-LLM Recommendation
di: Sanchez, Huascar, et al.
Pubblicazione: (2025)
di: Sanchez, Huascar, et al.
Pubblicazione: (2025)
Removing Spurious Correlation from Neural Network Interpretations
di: Fotouhi, Milad, et al.
Pubblicazione: (2024)
di: Fotouhi, Milad, et al.
Pubblicazione: (2024)
User-LLM: Efficient LLM Contextualization with User Embeddings
di: Ning, Lin, et al.
Pubblicazione: (2024)
di: Ning, Lin, et al.
Pubblicazione: (2024)
Reactive Transformer (RxT) -- Stateful Real-Time Processing for Event-Driven Reactive Language Models
di: Filipek, Adam
Pubblicazione: (2025)
di: Filipek, Adam
Pubblicazione: (2025)
LogProber: Disentangling confidence from contamination in LLM responses
di: Yax, Nicolas, et al.
Pubblicazione: (2024)
di: Yax, Nicolas, et al.
Pubblicazione: (2024)
VBART: The Turkish LLM
di: Turker, Meliksah, et al.
Pubblicazione: (2024)
di: Turker, Meliksah, et al.
Pubblicazione: (2024)
LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
di: Fu, Qichen, et al.
Pubblicazione: (2024)
di: Fu, Qichen, et al.
Pubblicazione: (2024)
BPO: Staying Close to the Behavior LLM Creates Better Online LLM Alignment
di: Xu, Wenda, et al.
Pubblicazione: (2024)
di: Xu, Wenda, et al.
Pubblicazione: (2024)
Large Language Models Assume People are More Rational than We Really are
di: Liu, Ryan, et al.
Pubblicazione: (2024)
di: Liu, Ryan, et al.
Pubblicazione: (2024)
LLM Maybe LongLM: Self-Extend LLM Context Window Without Tuning
di: Jin, Hongye, et al.
Pubblicazione: (2024)
di: Jin, Hongye, et al.
Pubblicazione: (2024)
LLM See, LLM Do: Guiding Data Generation to Target Non-Differentiable Objectives
di: Shimabucoro, Luísa, et al.
Pubblicazione: (2024)
di: Shimabucoro, Luísa, et al.
Pubblicazione: (2024)
MixEval: Deriving Wisdom of the Crowd from LLM Benchmark Mixtures
di: Ni, Jinjie, et al.
Pubblicazione: (2024)
di: Ni, Jinjie, et al.
Pubblicazione: (2024)
R-Zero: Self-Evolving Reasoning LLM from Zero Data
di: Huang, Chengsong, et al.
Pubblicazione: (2025)
di: Huang, Chengsong, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Harry Potter is Still Here! Probing Knowledge Leakage in Targeted Unlearned Large Language Models via Automated Adversarial Prompting
di: To, Bang Trinh Tran, et al.
Pubblicazione: (2025) -
Who can we trust? LLM-as-a-jury for Comparative Assessment
di: Qian, Mengjie, et al.
Pubblicazione: (2026) -
The illusion of a perfect metric: Why evaluating AI's words is harder than it looks
di: Oliva, Maria Paz, et al.
Pubblicazione: (2025) -
Who Gets Which Message? Auditing Demographic Bias in LLM-Generated Targeted Text
di: Islam, Tunazzina
Pubblicazione: (2026) -
Robustly Improving LLM Fairness in Realistic Settings via Interpretability
di: Karvonen, Adam, et al.
Pubblicazione: (2025)