Navigating the Alignment-Calibration Trade-off: A Pareto-Superior Frontier via Model Merging
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918180204576768 |
|---|---|
| author | Hu, Tiancheng Minixhofer, Benjamin Collier, Nigel |
| author_facet | Hu, Tiancheng Minixhofer, Benjamin Collier, Nigel |
| contents | The "alignment tax" of post-training is typically framed as a drop in task accuracy. We show it also involves a severe loss of calibration, making models overconfident, less reliable, and model outputs less diverse. We show that this trade-off can be navigated effectively via a simple post-hoc intervention: interpolating between a model's weights before and after alignment. Crucially, this is not a strict trade-off. We find that the process consistently reveals Pareto-optimal interpolations - models that improve accuracy beyond both parents while substantially recovering the calibration lost during alignment. Our work demonstrates that simple model merging provides a computationally efficient method for mitigating the full scope of the alignment tax, yielding models that are more capable and more reliable. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_17426 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Navigating the Alignment-Calibration Trade-off: A Pareto-Superior Frontier via Model Merging Hu, Tiancheng Minixhofer, Benjamin Collier, Nigel Computation and Language Artificial Intelligence Machine Learning The "alignment tax" of post-training is typically framed as a drop in task accuracy. We show it also involves a severe loss of calibration, making models overconfident, less reliable, and model outputs less diverse. We show that this trade-off can be navigated effectively via a simple post-hoc intervention: interpolating between a model's weights before and after alignment. Crucially, this is not a strict trade-off. We find that the process consistently reveals Pareto-optimal interpolations - models that improve accuracy beyond both parents while substantially recovering the calibration lost during alignment. Our work demonstrates that simple model merging provides a computationally efficient method for mitigating the full scope of the alignment tax, yielding models that are more capable and more reliable. |
| title | Navigating the Alignment-Calibration Trade-off: A Pareto-Superior Frontier via Model Merging |
| topic | Computation and Language Artificial Intelligence Machine Learning |
| url | https://arxiv.org/abs/2510.17426 |