Model Internal Sleuthing: Finding Lexical Identity and Inflectional Features in Modern Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Michael, Subramani, Nishant
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911613376790528
author Li, Michael
Subramani, Nishant
author_facet Li, Michael
Subramani, Nishant
contents Large transformer-based language models dominate modern NLP, yet our understanding of how they encode linguistic information relies primarily on studies of early models like BERT and GPT-2. We systematically probe 25 models from BERT Base to Qwen2.5-7B focusing on two linguistic properties: lexical identity and inflectional features across 6 diverse languages. We find a consistent pattern: inflectional features are linearly decodable throughout the model, while lexical identity is prominent early but increasingly weakens with depth. Further analysis of the representation geometry reveals that models with aggressive mid-layer dimensionality compression show reduced steering effectiveness in those layers, despite probe accuracy remaining high. Pretraining analysis shows that inflectional structure stabilizes early while lexical identity representations continue evolving. Taken together, our findings suggest that transformers maintain inflectional features across layers, while trading off lexical identity for compact, predictive representations. Our code is available at https://github.com/ml5885/model_internal_sleuthing
format Preprint
id arxiv_https___arxiv_org_abs_2506_02132
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Model Internal Sleuthing: Finding Lexical Identity and Inflectional Features in Modern Language Models
Li, Michael
Subramani, Nishant
Computation and Language
Machine Learning
Large transformer-based language models dominate modern NLP, yet our understanding of how they encode linguistic information relies primarily on studies of early models like BERT and GPT-2. We systematically probe 25 models from BERT Base to Qwen2.5-7B focusing on two linguistic properties: lexical identity and inflectional features across 6 diverse languages. We find a consistent pattern: inflectional features are linearly decodable throughout the model, while lexical identity is prominent early but increasingly weakens with depth. Further analysis of the representation geometry reveals that models with aggressive mid-layer dimensionality compression show reduced steering effectiveness in those layers, despite probe accuracy remaining high. Pretraining analysis shows that inflectional structure stabilizes early while lexical identity representations continue evolving. Taken together, our findings suggest that transformers maintain inflectional features across layers, while trading off lexical identity for compact, predictive representations. Our code is available at https://github.com/ml5885/model_internal_sleuthing
title Model Internal Sleuthing: Finding Lexical Identity and Inflectional Features in Modern Language Models
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2506.02132