A Comparative Analysis of Contextual Representation Flow in State-Space and Transformer Architectures

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Hoang, Nhat M., Long, Do Xuan, Nguyen, Cong-Duy, Kan, Min-Yen, Tuan, Luu Anh
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917187478880256
author Hoang, Nhat M.
Long, Do Xuan
Nguyen, Cong-Duy
Kan, Min-Yen
Tuan, Luu Anh
author_facet Hoang, Nhat M.
Long, Do Xuan
Nguyen, Cong-Duy
Kan, Min-Yen
Tuan, Luu Anh
contents State Space Models (SSMs) have recently emerged as efficient alternatives to Transformer-Based Models (TBMs) for long-sequence processing with linear scaling, yet how contextual information flows across layers in these architectures remains understudied. We present the first unified, token- and layer-wise analysis of representation propagation in SSMs and TBMs. Using centered kernel alignment, variance-based metrics, and probing, we characterize how representations evolve within and across layers. We find a key divergence: TBMs rapidly homogenize token representations, with diversity reemerging only in later layers, while SSMs preserve token uniqueness early but converge to homogenization deeper. Theoretical analysis and parameter randomization further reveal that oversmoothing in TBMs stems from architectural design, whereas in SSMs, it arises mainly from training dynamics. These insights clarify the inductive biases of both architectures and inform future model and training designs for long-context reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2510_06640
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Comparative Analysis of Contextual Representation Flow in State-Space and Transformer Architectures
Hoang, Nhat M.
Long, Do Xuan
Nguyen, Cong-Duy
Kan, Min-Yen
Tuan, Luu Anh
Computation and Language
Machine Learning
State Space Models (SSMs) have recently emerged as efficient alternatives to Transformer-Based Models (TBMs) for long-sequence processing with linear scaling, yet how contextual information flows across layers in these architectures remains understudied. We present the first unified, token- and layer-wise analysis of representation propagation in SSMs and TBMs. Using centered kernel alignment, variance-based metrics, and probing, we characterize how representations evolve within and across layers. We find a key divergence: TBMs rapidly homogenize token representations, with diversity reemerging only in later layers, while SSMs preserve token uniqueness early but converge to homogenization deeper. Theoretical analysis and parameter randomization further reveal that oversmoothing in TBMs stems from architectural design, whereas in SSMs, it arises mainly from training dynamics. These insights clarify the inductive biases of both architectures and inform future model and training designs for long-context reasoning.
title A Comparative Analysis of Contextual Representation Flow in State-Space and Transformer Architectures
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2510.06640