Knowledge Adaptation as Posterior Correction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Khan, Mohammad Emtiyaz
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908698860847104
author Khan, Mohammad Emtiyaz
author_facet Khan, Mohammad Emtiyaz
contents Adaptation is the holy grail of intelligence, but even the best AI models lack the adaptability of toddlers. In spite of great progress, little is known about the mechanisms by which machines can learn to adapt as fast as humans and animals. Here, we cast adaptation as `correction' of old posteriors and show that a wide-variety of existing adaptation methods follow this very principle, including those used for continual learning, federated learning, unlearning, and model merging. In all these settings, more accurate posteriors often lead to smaller corrections and can enable faster adaptation. Posterior correction is derived by using the dual representation of the Bayesian Learning Rule of Khan and Rue (2023), where the interference between the old representation and new information is quantified by using the natural-gradient mismatch. We present many examples demonstrating how machines can learn to adapt quickly by using posterior correction.
format Preprint
id arxiv_https___arxiv_org_abs_2506_14262
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Knowledge Adaptation as Posterior Correction
Khan, Mohammad Emtiyaz
Machine Learning
Artificial Intelligence
Adaptation is the holy grail of intelligence, but even the best AI models lack the adaptability of toddlers. In spite of great progress, little is known about the mechanisms by which machines can learn to adapt as fast as humans and animals. Here, we cast adaptation as `correction' of old posteriors and show that a wide-variety of existing adaptation methods follow this very principle, including those used for continual learning, federated learning, unlearning, and model merging. In all these settings, more accurate posteriors often lead to smaller corrections and can enable faster adaptation. Posterior correction is derived by using the dual representation of the Bayesian Learning Rule of Khan and Rue (2023), where the interference between the old representation and new information is quantified by using the natural-gradient mismatch. We present many examples demonstrating how machines can learn to adapt quickly by using posterior correction.
title Knowledge Adaptation as Posterior Correction
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2506.14262