Guardado en:
Detalles Bibliográficos
Autores principales: YILDIRIM, Alper, yucedag, ibrahim
Formato: Recurso digital
Lenguaje:
Publicado: Zenodo 2025
Materias:
Acceso en línea:https://doi.org/10.5281/zenodo.17667170
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866902247054508032
author YILDIRIM, Alper
yucedag, ibrahim
author_facet YILDIRIM, Alper
yucedag, ibrahim
contents <p><strong>Abstract</strong><br>This position paper argues that standard Transformers suffer from a fundamental structural inefficiency we term the “Semantic Alignment Tax.”<br>This tax represents the prohibitive optimization<br>cost of learning a coherent relational geometry<br>from a chaotic random initialization. To isolate<br>this phenomenon, we employ Iterative Semantic<br>Map Refinement (ISMR) as a diagnostic protocol.<br>By applying this probe to both deep and wide architectures, we uncover a critical invariance: the<br>alignment tax constitutes a fixed geometric barrier<br>that persists regardless of model depth or parameter ratio. Our findings validate the “Sculptor’s<br>Dilemma.” demonstrating that deep reasoning<br>layers cannot effectively refine representations<br>that are initially incoherent without expending<br>a significant portion of the compute budget on<br>geometric alignment. We contend that current<br>scaling strategies merely mask this tax rather than<br>solving it, a strategy that becomes untenable for<br>low-resource tasks. We call for a shift toward<br>“Representation-First” architectures that enforce<br>relational geometry at initialization.</p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_17667170
institution Zenodo
language
publishDate 2025
publisher Zenodo
record_format zenodo
spellingShingle Pay Attention Later: The Semantic Alignment Tax Requires Decoupling Representation from Reasoning
YILDIRIM, Alper
yucedag, ibrahim
Representation Learning
Transformer Optimization
Inductive Bias
Semantic Geometry
Initialization Schemes
Low-Resource NLP
Semantic Alignment Tax
Deep learning
Transformers
Neuro-AI
<p><strong>Abstract</strong><br>This position paper argues that standard Transformers suffer from a fundamental structural inefficiency we term the “Semantic Alignment Tax.”<br>This tax represents the prohibitive optimization<br>cost of learning a coherent relational geometry<br>from a chaotic random initialization. To isolate<br>this phenomenon, we employ Iterative Semantic<br>Map Refinement (ISMR) as a diagnostic protocol.<br>By applying this probe to both deep and wide architectures, we uncover a critical invariance: the<br>alignment tax constitutes a fixed geometric barrier<br>that persists regardless of model depth or parameter ratio. Our findings validate the “Sculptor’s<br>Dilemma.” demonstrating that deep reasoning<br>layers cannot effectively refine representations<br>that are initially incoherent without expending<br>a significant portion of the compute budget on<br>geometric alignment. We contend that current<br>scaling strategies merely mask this tax rather than<br>solving it, a strategy that becomes untenable for<br>low-resource tasks. We call for a shift toward<br>“Representation-First” architectures that enforce<br>relational geometry at initialization.</p>
title Pay Attention Later: The Semantic Alignment Tax Requires Decoupling Representation from Reasoning
topic Representation Learning
Transformer Optimization
Inductive Bias
Semantic Geometry
Initialization Schemes
Low-Resource NLP
Semantic Alignment Tax
Deep learning
Transformers
Neuro-AI
url https://doi.org/10.5281/zenodo.17667170