Provable unlearning in topic modeling and downstream tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Wei, Stanley, Malladi, Sadhika, Arora, Sanjeev, Sanyal, Amartya |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LESS: Selecting Influential Data for Targeted Instruction Tuning
by: Xia, Mengzhou, et al.
Published: (2024)
by: Xia, Mengzhou, et al.
Published: (2024)
Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization
by: Razin, Noam, et al.
Published: (2024)
by: Razin, Noam, et al.
Published: (2024)
Provable Privacy with Non-Private Pre-Processing
by: Hu, Yaxi, et al.
Published: (2024)
by: Hu, Yaxi, et al.
Published: (2024)
On the SDEs and Scaling Rules for Adaptive Gradient Algorithms
by: Malladi, Sadhika, et al.
Published: (2022)
by: Malladi, Sadhika, et al.
Published: (2022)
Trainable Transformer in Transformer
by: Panigrahi, Abhishek, et al.
Published: (2023)
by: Panigrahi, Abhishek, et al.
Published: (2023)
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient
by: Shang, Shuning, et al.
Published: (2026)
by: Shang, Shuning, et al.
Published: (2026)
Why pre-training is beneficial for downstream classification tasks?
by: Jiang, Xin, et al.
Published: (2024)
by: Jiang, Xin, et al.
Published: (2024)
Applying sparse autoencoders to unlearn knowledge in language models
by: Farrell, Eoin, et al.
Published: (2024)
by: Farrell, Eoin, et al.
Published: (2024)
Machine unlearning through fine-grained model parameters perturbation
by: Zuo, Zhiwei, et al.
Published: (2024)
by: Zuo, Zhiwei, et al.
Published: (2024)
Training data attribution in diffusion models via mirrored unlearning and noise-consistent skew
by: Serrà, Joan, et al.
Published: (2026)
by: Serrà, Joan, et al.
Published: (2026)
Preference Learning Algorithms Do Not Learn Preference Rankings
by: Chen, Angelica, et al.
Published: (2024)
by: Chen, Angelica, et al.
Published: (2024)
What Makes a Reward Model a Good Teacher? An Optimization Perspective
by: Razin, Noam, et al.
Published: (2025)
by: Razin, Noam, et al.
Published: (2025)
Accuracy on the wrong line: On the pitfalls of noisy data for out-of-distribution generalisation
by: Sanyal, Amartya, et al.
Published: (2024)
by: Sanyal, Amartya, et al.
Published: (2024)
On the limitation of evaluating machine unlearning using only a single training seed
by: Lanyon, Jamie, et al.
Published: (2025)
by: Lanyon, Jamie, et al.
Published: (2025)
Skill-Targeted Adaptive Training
by: He, Yinghui, et al.
Published: (2025)
by: He, Yinghui, et al.
Published: (2025)
Metacognitive Reuse: Turning Recurring LLM Reasoning Into Concise Behaviors
by: Didolkar, Aniket, et al.
Published: (2025)
by: Didolkar, Aniket, et al.
Published: (2025)
The Coverage Principle: How Pre-Training Enables Post-Training
by: Chen, Fan, et al.
Published: (2025)
by: Chen, Fan, et al.
Published: (2025)
Modular Jets for Supervised Pipelines: Diagnosing Mirage vs Identifiability
by: Sanyal, Suman
Published: (2025)
by: Sanyal, Suman
Published: (2025)
Perception Learning: A Formal Separation of Sensory Representation Learning from Decision Learning
by: Sanyal, Suman
Published: (2025)
by: Sanyal, Suman
Published: (2025)
Self-Play with Adversarial Critic: Provable and Scalable Offline Alignment for Language Models
by: Ji, Xiang, et al.
Published: (2024)
by: Ji, Xiang, et al.
Published: (2024)
Fine-Tuning Language Models with Just Forward Passes
by: Malladi, Sadhika, et al.
Published: (2023)
by: Malladi, Sadhika, et al.
Published: (2023)
Why is Your Language Model a Poor Implicit Reward Model?
by: Razin, Noam, et al.
Published: (2025)
by: Razin, Noam, et al.
Published: (2025)
Contextual Drag: How Errors in the Context Affect LLM Reasoning
by: Cheng, Yun, et al.
Published: (2026)
by: Cheng, Yun, et al.
Published: (2026)
Unrealized Expectations: Comparing AI Methods vs Classical Algorithms for Maximum Independent Set
by: Wu, Yikai, et al.
Published: (2025)
by: Wu, Yikai, et al.
Published: (2025)
Provable Training Data Identification for Large Language Models
by: Liu, Zhenlong, et al.
Published: (2025)
by: Liu, Zhenlong, et al.
Published: (2025)
Corrective Machine Unlearning
by: Goel, Shashwat, et al.
Published: (2024)
by: Goel, Shashwat, et al.
Published: (2024)
HiVAE: Hierarchical Latent Variables for Scalable Theory of Mind
by: Doering, Nigel, et al.
Published: (2026)
by: Doering, Nigel, et al.
Published: (2026)
Model-agnostic Selective Labeling with Provable Statistical Guarantees
by: Huang, Huipeng, et al.
Published: (2025)
by: Huang, Huipeng, et al.
Published: (2025)
Can Models Learn Skill Composition from Examples?
by: Zhao, Haoyu, et al.
Published: (2024)
by: Zhao, Haoyu, et al.
Published: (2024)
Provable In-Context Vector Arithmetic via Retrieving Task Concepts
by: Bu, Dake, et al.
Published: (2025)
by: Bu, Dake, et al.
Published: (2025)
Unlearning via Sparse Representations
by: Shah, Vedant, et al.
Published: (2023)
by: Shah, Vedant, et al.
Published: (2023)
Provably Learning Diffusion Models under the Manifold Hypothesis: Collapse and Refine
by: Huang, Wei, et al.
Published: (2026)
by: Huang, Wei, et al.
Published: (2026)
Provable Robust Saliency-based Explanations
by: Chen, Chao, et al.
Published: (2022)
by: Chen, Chao, et al.
Published: (2022)
Provable Generalization in Overparameterized Neural Nets
by: Dhingra, Aviral
Published: (2025)
by: Dhingra, Aviral
Published: (2025)
Conditional Equivalence of DPO and RLHF: Implicit Assumption, Failure Modes, and Provable Alignment
by: Yang, Zhiqin, et al.
Published: (2026)
by: Yang, Zhiqin, et al.
Published: (2026)
Second-Order Actor-Critic Methods for Discounted MDPs via Policy Hessian Decomposition
by: Manivannan, Sanjeev, et al.
Published: (2026)
by: Manivannan, Sanjeev, et al.
Published: (2026)
Dichotomy of Early and Late Phase Implicit Biases Can Provably Induce Grokking
by: Lyu, Kaifeng, et al.
Published: (2023)
by: Lyu, Kaifeng, et al.
Published: (2023)
An empirical study of task and feature correlations in the reuse of pre-trained models
by: Mohamud, Jama Hussein, et al.
Published: (2025)
by: Mohamud, Jama Hussein, et al.
Published: (2025)
Keeping LLMs Aligned After Fine-tuning: The Crucial Role of Prompt Templates
by: Lyu, Kaifeng, et al.
Published: (2024)
by: Lyu, Kaifeng, et al.
Published: (2024)
Guide: Generalized-Prior and Data Encoders for DAG Estimation
by: Roy, Amartya, et al.
Published: (2025)
by: Roy, Amartya, et al.
Published: (2025)
Similar Items
-
LESS: Selecting Influential Data for Targeted Instruction Tuning
by: Xia, Mengzhou, et al.
Published: (2024) -
Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization
by: Razin, Noam, et al.
Published: (2024) -
Provable Privacy with Non-Private Pre-Processing
by: Hu, Yaxi, et al.
Published: (2024) -
On the SDEs and Scaling Rules for Adaptive Gradient Algorithms
by: Malladi, Sadhika, et al.
Published: (2022) -
Trainable Transformer in Transformer
by: Panigrahi, Abhishek, et al.
Published: (2023)