Beyond Temperature: Hyperfitting as a Late-Stage Geometric Expansion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Meimingwei, Ding, Yuanhao, Arias, Esteban Garces, Heumann, Christian
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913152949551104
author Li, Meimingwei
Ding, Yuanhao
Arias, Esteban Garces
Heumann, Christian
author_facet Li, Meimingwei
Ding, Yuanhao
Arias, Esteban Garces
Heumann, Christian
contents Recent work has identified a counterintuitive phenomenon termed "Hyperfitting", where fine-tuning Large Language Models (LLMs) to near-zero training loss on small datasets surprisingly enhances open-ended generation quality and mitigates repetition in greedy decoding. While effective, the underlying mechanism remains poorly understood, with the extremely low-entropy output distributions suggesting a potential equivalence to simple temperature scaling. In this work, we demonstrate that this phenomenon is fundamentally distinct from distribution sharpening; entropy-matched control experiments reveal that temperature scaling fails to replicate the diversity gains of hyperfitting. Furthermore, we falsify the hypothesis of static vocabulary reweighting, showing through ablation studies that hyperfitting relies on a dynamic, context-dependent rank reordering mechanism. Layer-wise analysis localizes this effect to a "Terminal Expansion" in the final transformer block, where a substantial geometric expansion of the feature space (Delta Dim approx +80.8) facilitates the promotion of deep-tail tokens. Additionally, we introduce Late-Stage LoRA, a targeted fine-tuning strategy that updates only the final 5 layers, yielding robust generation with minimal parameter updates
format Preprint
id arxiv_https___arxiv_org_abs_2605_22579
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Beyond Temperature: Hyperfitting as a Late-Stage Geometric Expansion
Li, Meimingwei
Ding, Yuanhao
Arias, Esteban Garces
Heumann, Christian
Computation and Language
Artificial Intelligence
Machine Learning
Recent work has identified a counterintuitive phenomenon termed "Hyperfitting", where fine-tuning Large Language Models (LLMs) to near-zero training loss on small datasets surprisingly enhances open-ended generation quality and mitigates repetition in greedy decoding. While effective, the underlying mechanism remains poorly understood, with the extremely low-entropy output distributions suggesting a potential equivalence to simple temperature scaling. In this work, we demonstrate that this phenomenon is fundamentally distinct from distribution sharpening; entropy-matched control experiments reveal that temperature scaling fails to replicate the diversity gains of hyperfitting. Furthermore, we falsify the hypothesis of static vocabulary reweighting, showing through ablation studies that hyperfitting relies on a dynamic, context-dependent rank reordering mechanism. Layer-wise analysis localizes this effect to a "Terminal Expansion" in the final transformer block, where a substantial geometric expansion of the feature space (Delta Dim approx +80.8) facilitates the promotion of deep-tail tokens. Additionally, we introduce Late-Stage LoRA, a targeted fine-tuning strategy that updates only the final 5 layers, yielding robust generation with minimal parameter updates
title Beyond Temperature: Hyperfitting as a Late-Stage Geometric Expansion
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2605.22579