Saved in:
| Main Authors: | McKinney, Lev, Thudi, Anvith, Bae, Juhan, Rezaei, Tara, Papernot, Nicolas, McIlraith, Sheila A., Grosse, Roger |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2602.10568 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Better Training Data Attribution via Better Inverse Hessian-Vector Products
by: Wang, Andrew, et al.
Published: (2025)
by: Wang, Andrew, et al.
Published: (2025)
Fast Exact Unlearning for In-Context Learning Data for LLMs
by: Muresanu, Andrei I., et al.
Published: (2024)
by: Muresanu, Andrei I., et al.
Published: (2024)
Gradients Look Alike: Sensitivity is Often Overestimated in DP-SGD
by: Thudi, Anvith, et al.
Published: (2023)
by: Thudi, Anvith, et al.
Published: (2023)
Efficient Public Verification of Private ML via Regularization
by: Bell, Zoë Ruha, et al.
Published: (2025)
by: Bell, Zoë Ruha, et al.
Published: (2025)
Leveraging Per-Instance Privacy for Machine Unlearning
by: Sepahvand, Nazanin Mohammadi, et al.
Published: (2025)
by: Sepahvand, Nazanin Mohammadi, et al.
Published: (2025)
MixMax: Distributional Robustness in Function Space via Optimal Data Mixtures
by: Thudi, Anvith, et al.
Published: (2024)
by: Thudi, Anvith, et al.
Published: (2024)
Pluralistic Alignment Over Time
by: Klassen, Toryn Q., et al.
Published: (2024)
by: Klassen, Toryn Q., et al.
Published: (2024)
Formal Methods Meet LLMs: Auditing, Monitoring, and Intervention for Compliance of Advanced AI Systems
by: Alamdari, Parand A., et al.
Published: (2026)
by: Alamdari, Parand A., et al.
Published: (2026)
STEVE-1: A Generative Model for Text-to-Behavior in Minecraft
by: Lifshitz, Shalev, et al.
Published: (2023)
by: Lifshitz, Shalev, et al.
Published: (2023)
Remembering to Be Fair: Non-Markovian Fairness in Sequential Decision Making
by: Alamdari, Parand A., et al.
Published: (2023)
by: Alamdari, Parand A., et al.
Published: (2023)
Selective Prediction via Training Dynamics
by: Rabanser, Stephan, et al.
Published: (2022)
by: Rabanser, Stephan, et al.
Published: (2022)
Being Considerate as a Pathway Towards Pluralistic Alignment for Agentic AI
by: Alamdari, Parand A., et al.
Published: (2024)
by: Alamdari, Parand A., et al.
Published: (2024)
Pushdown Reward Machines for Reinforcement Learning
by: Varricchione, Giovanni, et al.
Published: (2025)
by: Varricchione, Giovanni, et al.
Published: (2025)
Ground-Compose-Reinforce: Grounding Language in Agentic Behaviours using Limited Data
by: Li, Andrew C., et al.
Published: (2025)
by: Li, Andrew C., et al.
Published: (2025)
MixMin: Finding Data Mixtures via Convex Minimization
by: Thudi, Anvith, et al.
Published: (2025)
by: Thudi, Anvith, et al.
Published: (2025)
Training Data Attribution via Approximate Unrolled Differentiation
by: Bae, Juhan, et al.
Published: (2024)
by: Bae, Juhan, et al.
Published: (2024)
Multi-Agent Verification: Scaling Test-Time Compute with Multiple Verifiers
by: Lifshitz, Shalev, et al.
Published: (2025)
by: Lifshitz, Shalev, et al.
Published: (2025)
Leakage Safe Graph Features for Interpretable Fraud Detection in Temporal Transaction Networks
by: Khaleghpour, Hamideh, et al.
Published: (2026)
by: Khaleghpour, Hamideh, et al.
Published: (2026)
Optimizing Neuro-Fuzzy and Colonial Competition Algorithms for Skin Cancer Diagnosis in Dermatoscopic Images
by: Khaleghpour, Hamideh, et al.
Published: (2025)
by: Khaleghpour, Hamideh, et al.
Published: (2025)
Unified AI for Accurate Audio Anomaly Detection
by: Khaleghpour, Hamideh, et al.
Published: (2025)
by: Khaleghpour, Hamideh, et al.
Published: (2025)
Now, Later, and Lasting: Ten Priorities for AI Research, Policy, and Practice
by: Horvitz, Eric, et al.
Published: (2024)
by: Horvitz, Eric, et al.
Published: (2024)
Reward Machines for Deep RL in Noisy and Uncertain Environments
by: Li, Andrew C., et al.
Published: (2024)
by: Li, Andrew C., et al.
Published: (2024)
Negation Neglect: When models fail to learn negations in training
by: Mayne, Harry, et al.
Published: (2026)
by: Mayne, Harry, et al.
Published: (2026)
Verifiable and Provably Secure Machine Unlearning
by: Eisenhofer, Thorsten, et al.
Published: (2022)
by: Eisenhofer, Thorsten, et al.
Published: (2022)
Error whitening: Why Gauss-Newton outperforms Newton
by: McKay, Maricela Best, et al.
Published: (2026)
by: McKay, Maricela Best, et al.
Published: (2026)
Language Models For Generalised PDDL Planning: Synthesising Sound and Programmatic Policies
by: Chen, Dillon Z., et al.
Published: (2025)
by: Chen, Dillon Z., et al.
Published: (2025)
Eliciting Latent Predictions from Transformers with the Tuned Lens
by: Belrose, Nora, et al.
Published: (2023)
by: Belrose, Nora, et al.
Published: (2023)
GULPS: Two-Qubit Gate Synthesis via Linear Programming for Heterogeneous Instruction Sets
by: McKinney, Evan, et al.
Published: (2025)
by: McKinney, Evan, et al.
Published: (2025)
Spectral-factorized Positive-definite Curvature Learning for NN Training
by: Lin, Wu, et al.
Published: (2025)
by: Lin, Wu, et al.
Published: (2025)
Inexact Unlearning Needs More Careful Evaluations to Avoid a False Sense of Privacy
by: Hayes, Jamie, et al.
Published: (2024)
by: Hayes, Jamie, et al.
Published: (2024)
Satisficing and Optimal Generalised Planning via Goal Regression (Extended Version)
by: Chen, Dillon Z., et al.
Published: (2025)
by: Chen, Dillon Z., et al.
Published: (2025)
Learning Bilevel Policies over Symbolic World Models for Long-Horizon Planning
by: Chen, Dillon Z., et al.
Published: (2026)
by: Chen, Dillon Z., et al.
Published: (2026)
What Does It Take to Build a Performant Selective Classifier?
by: Rabanser, Stephan, et al.
Published: (2025)
by: Rabanser, Stephan, et al.
Published: (2025)
Demystify Protein Generation with Hierarchical Conditional Diffusion Models
by: Ling, Zinan, et al.
Published: (2025)
by: Ling, Zinan, et al.
Published: (2025)
AutoPrognosis 2.0: Democratizing Diagnostic and Prognostic Modeling in Healthcare with Automated Machine Learning
by: Imrie, Fergus, et al.
Published: (2022)
by: Imrie, Fergus, et al.
Published: (2022)
LLM Dataset Inference: Did you train on my dataset?
by: Maini, Pratyush, et al.
Published: (2024)
by: Maini, Pratyush, et al.
Published: (2024)
Fairness Feedback Loops: Training on Synthetic Data Amplifies Bias
by: Wyllie, Sierra, et al.
Published: (2024)
by: Wyllie, Sierra, et al.
Published: (2024)
Concept Influence: Leveraging Interpretability to Improve Performance and Efficiency in Training Data Attribution
by: Kowal, Matthew, et al.
Published: (2026)
by: Kowal, Matthew, et al.
Published: (2026)
UnUnlearning: Unlearning is not sufficient for content regulation in advanced generative AI
by: Shumailov, Ilia, et al.
Published: (2024)
by: Shumailov, Ilia, et al.
Published: (2024)
Tighter Privacy Auditing of DP-SGD in the Hidden State Threat Model
by: Cebere, Tudor, et al.
Published: (2024)
by: Cebere, Tudor, et al.
Published: (2024)
Similar Items
-
Better Training Data Attribution via Better Inverse Hessian-Vector Products
by: Wang, Andrew, et al.
Published: (2025) -
Fast Exact Unlearning for In-Context Learning Data for LLMs
by: Muresanu, Andrei I., et al.
Published: (2024) -
Gradients Look Alike: Sensitivity is Often Overestimated in DP-SGD
by: Thudi, Anvith, et al.
Published: (2023) -
Efficient Public Verification of Private ML via Regularization
by: Bell, Zoë Ruha, et al.
Published: (2025) -
Leveraging Per-Instance Privacy for Machine Unlearning
by: Sepahvand, Nazanin Mohammadi, et al.
Published: (2025)