Less is More: Local Intrinsic Dimensions of Contextual Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ruppik, Benjamin Matthias, von Rohrscheidt, Julius, van Niekerk, Carel, Heck, Michael, Vukovic, Renato, Feng, Shutong, Lin, Hsien-chin, Lubis, Nurul, Rieck, Bastian, Zibrowius, Marcus, Gašić, Milica
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917044032634880
author Ruppik, Benjamin Matthias
von Rohrscheidt, Julius
van Niekerk, Carel
Heck, Michael
Vukovic, Renato
Feng, Shutong
Lin, Hsien-chin
Lubis, Nurul
Rieck, Bastian
Zibrowius, Marcus
Gašić, Milica
author_facet Ruppik, Benjamin Matthias
von Rohrscheidt, Julius
van Niekerk, Carel
Heck, Michael
Vukovic, Renato
Feng, Shutong
Lin, Hsien-chin
Lubis, Nurul
Rieck, Bastian
Zibrowius, Marcus
Gašić, Milica
contents Understanding the internal mechanisms of large language models (LLMs) remains a challenging and complex endeavor. Even fundamental questions, such as how fine-tuning affects model behavior, often require extensive empirical evaluation. In this paper, we introduce a novel perspective based on the geometric properties of contextual latent embeddings to study the effects of training and fine-tuning. To that end, we measure the local dimensions of a contextual language model's latent space and analyze their shifts during training and fine-tuning. We show that the local dimensions provide insights into the model's training dynamics and generalization ability. Specifically, the mean of the local dimensions predicts when the model's training capabilities are exhausted, as exemplified in a dialogue state tracking task, overfitting, as demonstrated in an emotion recognition task, and grokking, as illustrated with an arithmetic task. Furthermore, our experiments suggest a practical heuristic: reductions in the mean local dimension tend to accompany and predict subsequent performance gains. Through this exploration, we aim to provide practitioners with a deeper understanding of the implications of fine-tuning on embedding spaces, facilitating informed decisions when configuring models for specific applications. The results of this work contribute to the ongoing discourse on the interpretability, adaptability, and generalizability of LLMs by bridging the gap between intrinsic model mechanisms and geometric properties in the respective embeddings.
format Preprint
id arxiv_https___arxiv_org_abs_2506_01034
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Less is More: Local Intrinsic Dimensions of Contextual Language Models
Ruppik, Benjamin Matthias
von Rohrscheidt, Julius
van Niekerk, Carel
Heck, Michael
Vukovic, Renato
Feng, Shutong
Lin, Hsien-chin
Lubis, Nurul
Rieck, Bastian
Zibrowius, Marcus
Gašić, Milica
Computation and Language
Artificial Intelligence
Machine Learning
Understanding the internal mechanisms of large language models (LLMs) remains a challenging and complex endeavor. Even fundamental questions, such as how fine-tuning affects model behavior, often require extensive empirical evaluation. In this paper, we introduce a novel perspective based on the geometric properties of contextual latent embeddings to study the effects of training and fine-tuning. To that end, we measure the local dimensions of a contextual language model's latent space and analyze their shifts during training and fine-tuning. We show that the local dimensions provide insights into the model's training dynamics and generalization ability. Specifically, the mean of the local dimensions predicts when the model's training capabilities are exhausted, as exemplified in a dialogue state tracking task, overfitting, as demonstrated in an emotion recognition task, and grokking, as illustrated with an arithmetic task. Furthermore, our experiments suggest a practical heuristic: reductions in the mean local dimension tend to accompany and predict subsequent performance gains. Through this exploration, we aim to provide practitioners with a deeper understanding of the implications of fine-tuning on embedding spaces, facilitating informed decisions when configuring models for specific applications. The results of this work contribute to the ongoing discourse on the interpretability, adaptability, and generalizability of LLMs by bridging the gap between intrinsic model mechanisms and geometric properties in the respective embeddings.
title Less is More: Local Intrinsic Dimensions of Contextual Language Models
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2506.01034