TOFU: A Task of Fictitious Unlearning for LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Maini, Pratyush, Feng, Zhili, Schwarzschild, Avi, Lipton, Zachary C., Kolter, J. Zico |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rethinking LLM Memorization through the Lens of Adversarial Compression
by: Schwarzschild, Avi, et al.
Published: (2024)
by: Schwarzschild, Avi, et al.
Published: (2024)
T-MARS: Improving Visual Representations by Circumventing Text Feature Learning
by: Maini, Pratyush, et al.
Published: (2023)
by: Maini, Pratyush, et al.
Published: (2023)
Existing Large Language Model Unlearning Evaluations Are Inconclusive
by: Feng, Zhili, et al.
Published: (2025)
by: Feng, Zhili, et al.
Published: (2025)
Forcing Diffuse Distributions out of Language Models
by: Zhang, Yiming, et al.
Published: (2024)
by: Zhang, Yiming, et al.
Published: (2024)
Understanding Hallucinations in Diffusion Models through Mode Interpolation
by: Aithal, Sumukh K, et al.
Published: (2024)
by: Aithal, Sumukh K, et al.
Published: (2024)
OpenUnlearning: Accelerating LLM Unlearning via Unified Benchmarking of Methods and Metrics
by: Dorna, Vineeth, et al.
Published: (2025)
by: Dorna, Vineeth, et al.
Published: (2025)
Scaling Laws for Data Filtering -- Data Curation cannot be Compute Agnostic
by: Goyal, Sachin, et al.
Published: (2024)
by: Goyal, Sachin, et al.
Published: (2024)
Prompt Recovery for Image Generation Models: A Comparative Study of Discrete Optimizers
by: Williams, Joshua Nathaniel, et al.
Published: (2024)
by: Williams, Joshua Nathaniel, et al.
Published: (2024)
Predicting the Performance of Black-box LLMs through Follow-up Queries
by: Sam, Dylan, et al.
Published: (2025)
by: Sam, Dylan, et al.
Published: (2025)
FUSE-ing Language Models: Zero-Shot Adapter Discovery for Prompt Optimization Across Tokenizers
by: Williams, Joshua Nathaniel, et al.
Published: (2024)
by: Williams, Joshua Nathaniel, et al.
Published: (2024)
Peeking Behind Closed Doors: Risks of LLM Evaluation by Private Data Curators
by: Bansal, Hritik, et al.
Published: (2025)
by: Bansal, Hritik, et al.
Published: (2025)
When Should We Introduce Safety Interventions During Pretraining?
by: Sam, Dylan, et al.
Published: (2026)
by: Sam, Dylan, et al.
Published: (2026)
Context-Parametric Inversion: Why Instruction Finetuning Can Worsen Context Reliance
by: Goyal, Sachin, et al.
Published: (2024)
by: Goyal, Sachin, et al.
Published: (2024)
Massive Activations in Large Language Models
by: Sun, Mingjie, et al.
Published: (2024)
by: Sun, Mingjie, et al.
Published: (2024)
Antidistillation Sampling
by: Savani, Yash, et al.
Published: (2025)
by: Savani, Yash, et al.
Published: (2025)
Benchmarking ChatGPT on Algorithmic Reasoning
by: McLeish, Sean, et al.
Published: (2024)
by: McLeish, Sean, et al.
Published: (2024)
A Simple and Effective Pruning Approach for Large Language Models
by: Sun, Mingjie, et al.
Published: (2023)
by: Sun, Mingjie, et al.
Published: (2023)
STAMP Your Content: Proving Dataset Membership via Watermarked Rephrasings
by: Rastogi, Saksham, et al.
Published: (2025)
by: Rastogi, Saksham, et al.
Published: (2025)
Mimetic Initialization Helps State Space Models Learn to Recall
by: Trockman, Asher, et al.
Published: (2024)
by: Trockman, Asher, et al.
Published: (2024)
Looking beyond the next token
by: Thankaraj, Abitha, et al.
Published: (2025)
by: Thankaraj, Abitha, et al.
Published: (2025)
Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning
by: Xu, Yixuan Even, et al.
Published: (2025)
by: Xu, Yixuan Even, et al.
Published: (2025)
Failure Modes of LLMs for Causal Reasoning on Narratives
by: Yamin, Khurram, et al.
Published: (2024)
by: Yamin, Khurram, et al.
Published: (2024)
LLM Dataset Inference: Did you train on my dataset?
by: Maini, Pratyush, et al.
Published: (2024)
by: Maini, Pratyush, et al.
Published: (2024)
Toward Understanding the Transferability of Adversarial Suffixes in Large Language Models
by: Ball, Sarah, et al.
Published: (2025)
by: Ball, Sarah, et al.
Published: (2025)
R-TOFU: Unlearning in Large Reasoning Models
by: Yoon, Sangyeon, et al.
Published: (2025)
by: Yoon, Sangyeon, et al.
Published: (2025)
Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated Text
by: Hans, Abhimanyu, et al.
Published: (2024)
by: Hans, Abhimanyu, et al.
Published: (2024)
Base Models Look Human To AI Detectors
by: Xu, Yixuan Even, et al.
Published: (2026)
by: Xu, Yixuan Even, et al.
Published: (2026)
An Axiomatic Approach to Model-Agnostic Concept Explanations
by: Feng, Zhili, et al.
Published: (2024)
by: Feng, Zhili, et al.
Published: (2024)
Safety Pretraining: Toward the Next Generation of Safe AI
by: Maini, Pratyush, et al.
Published: (2025)
by: Maini, Pratyush, et al.
Published: (2025)
AcceleratedLiNGAM: Learning Causal DAGs at the speed of GPUs
by: Akinwande, Victor, et al.
Published: (2024)
by: Akinwande, Victor, et al.
Published: (2024)
Adaptive Data Optimization: Dynamic Sample Selection with Scaling Laws
by: Jiang, Yiding, et al.
Published: (2024)
by: Jiang, Yiding, et al.
Published: (2024)
Leverage Unlearning to Sanitize LLMs
by: Boutet, Antoine, et al.
Published: (2025)
by: Boutet, Antoine, et al.
Published: (2025)
Mimetic Initialization of MLPs
by: Trockman, Asher, et al.
Published: (2026)
by: Trockman, Asher, et al.
Published: (2026)
LLM-Select: Feature Selection with Large Language Models
by: Jeong, Daniel P., et al.
Published: (2024)
by: Jeong, Daniel P., et al.
Published: (2024)
GRU: Mitigating the Trade-off between Unlearning and Retention for LLMs
by: Wang, Yue, et al.
Published: (2025)
by: Wang, Yue, et al.
Published: (2025)
Training a Generally Curious Agent
by: Tajwar, Fahim, et al.
Published: (2025)
by: Tajwar, Fahim, et al.
Published: (2025)
Personalized Language Modeling from Personalized Human Feedback
by: Li, Xinyu, et al.
Published: (2024)
by: Li, Xinyu, et al.
Published: (2024)
Evaluating the Factuality of Zero-shot Summarizers Across Varied Domains
by: Ramprasad, Sanjana, et al.
Published: (2024)
by: Ramprasad, Sanjana, et al.
Published: (2024)
Learn and Unlearn: Addressing Misinformation in Multilingual LLMs
by: Lu, Taiming, et al.
Published: (2024)
by: Lu, Taiming, et al.
Published: (2024)
Learning-Time Encoding Shapes Unlearning in LLMs
by: Wu, Ruihan, et al.
Published: (2025)
by: Wu, Ruihan, et al.
Published: (2025)
Similar Items
-
Rethinking LLM Memorization through the Lens of Adversarial Compression
by: Schwarzschild, Avi, et al.
Published: (2024) -
T-MARS: Improving Visual Representations by Circumventing Text Feature Learning
by: Maini, Pratyush, et al.
Published: (2023) -
Existing Large Language Model Unlearning Evaluations Are Inconclusive
by: Feng, Zhili, et al.
Published: (2025) -
Forcing Diffuse Distributions out of Language Models
by: Zhang, Yiming, et al.
Published: (2024) -
Understanding Hallucinations in Diffusion Models through Mode Interpolation
by: Aithal, Sumukh K, et al.
Published: (2024)