TIDE: Every Layer Knows the Token Beneath the Context

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jaiswal, Ajay, Hannah, Lauren, Kim, Han-Byul, Hoang, Duc, Farajtabar, Mehrdad, Cho, Minsik
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911657353019392
author Jaiswal, Ajay
Hannah, Lauren
Kim, Han-Byul
Hoang, Duc
Farajtabar, Mehrdad
Cho, Minsik
author_facet Jaiswal, Ajay
Hannah, Lauren
Kim, Han-Byul
Hoang, Duc
Farajtabar, Mehrdad
Cho, Minsik
contents We revisit a universally accepted but under-examined design choice in every modern LLM: a token index is looked up once at the input embedding layer and then permanently discarded. This single-injection assumption induces two structural failures: (i) the Rare Token Problem, where a Zipf-type distribution of vocabulary causes rare-token embeddings are chronically under-trained due to receiving a fraction of the cumulative gradient signal compared to common tokens; and (ii) the Contextual Collapse Problem, where limited parameters models map distributionally similar tokens to indistinguishable hidden states. As an attempt to address both, we propose TIDE, which augments the standard transformer with EmbeddingMemory: an ensemble of K independent MemoryBlocks that map token indices to context-free semantic vectors, computed once and injected into every layer through a depth-conditioned softmax router with a learnable null bank. We theoretically and empirically establish the benefits of TIDE in addressing the issues associated with single-token identity injection as well as improve performance across multiple language modeling and downstream tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2605_06216
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle TIDE: Every Layer Knows the Token Beneath the Context
Jaiswal, Ajay
Hannah, Lauren
Kim, Han-Byul
Hoang, Duc
Farajtabar, Mehrdad
Cho, Minsik
Computation and Language
Artificial Intelligence
Machine Learning
We revisit a universally accepted but under-examined design choice in every modern LLM: a token index is looked up once at the input embedding layer and then permanently discarded. This single-injection assumption induces two structural failures: (i) the Rare Token Problem, where a Zipf-type distribution of vocabulary causes rare-token embeddings are chronically under-trained due to receiving a fraction of the cumulative gradient signal compared to common tokens; and (ii) the Contextual Collapse Problem, where limited parameters models map distributionally similar tokens to indistinguishable hidden states. As an attempt to address both, we propose TIDE, which augments the standard transformer with EmbeddingMemory: an ensemble of K independent MemoryBlocks that map token indices to context-free semantic vectors, computed once and injected into every layer through a depth-conditioned softmax router with a learnable null bank. We theoretically and empirically establish the benefits of TIDE in addressing the issues associated with single-token identity injection as well as improve performance across multiple language modeling and downstream tasks.
title TIDE: Every Layer Knows the Token Beneath the Context
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2605.06216