Guardado en:
| Autores principales: | , |
|---|---|
| Formato: | Recurso digital |
| Lenguaje: | |
| Publicado: |
Zenodo
2025
|
| Materias: | |
| Acceso en línea: | https://doi.org/10.5281/zenodo.17667170 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866902247054508032 |
|---|---|
| author | YILDIRIM, Alper yucedag, ibrahim |
| author_facet | YILDIRIM, Alper yucedag, ibrahim |
| contents | <p><strong>Abstract</strong><br>This position paper argues that standard Transformers suffer from a fundamental structural inefficiency we term the “Semantic Alignment Tax.”<br>This tax represents the prohibitive optimization<br>cost of learning a coherent relational geometry<br>from a chaotic random initialization. To isolate<br>this phenomenon, we employ Iterative Semantic<br>Map Refinement (ISMR) as a diagnostic protocol.<br>By applying this probe to both deep and wide architectures, we uncover a critical invariance: the<br>alignment tax constitutes a fixed geometric barrier<br>that persists regardless of model depth or parameter ratio. Our findings validate the “Sculptor’s<br>Dilemma.” demonstrating that deep reasoning<br>layers cannot effectively refine representations<br>that are initially incoherent without expending<br>a significant portion of the compute budget on<br>geometric alignment. We contend that current<br>scaling strategies merely mask this tax rather than<br>solving it, a strategy that becomes untenable for<br>low-resource tasks. We call for a shift toward<br>“Representation-First” architectures that enforce<br>relational geometry at initialization.</p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_17667170 |
| institution | Zenodo |
| language | |
| publishDate | 2025 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | Pay Attention Later: The Semantic Alignment Tax Requires Decoupling Representation from Reasoning YILDIRIM, Alper yucedag, ibrahim Representation Learning Transformer Optimization Inductive Bias Semantic Geometry Initialization Schemes Low-Resource NLP Semantic Alignment Tax Deep learning Transformers Neuro-AI <p><strong>Abstract</strong><br>This position paper argues that standard Transformers suffer from a fundamental structural inefficiency we term the “Semantic Alignment Tax.”<br>This tax represents the prohibitive optimization<br>cost of learning a coherent relational geometry<br>from a chaotic random initialization. To isolate<br>this phenomenon, we employ Iterative Semantic<br>Map Refinement (ISMR) as a diagnostic protocol.<br>By applying this probe to both deep and wide architectures, we uncover a critical invariance: the<br>alignment tax constitutes a fixed geometric barrier<br>that persists regardless of model depth or parameter ratio. Our findings validate the “Sculptor’s<br>Dilemma.” demonstrating that deep reasoning<br>layers cannot effectively refine representations<br>that are initially incoherent without expending<br>a significant portion of the compute budget on<br>geometric alignment. We contend that current<br>scaling strategies merely mask this tax rather than<br>solving it, a strategy that becomes untenable for<br>low-resource tasks. We call for a shift toward<br>“Representation-First” architectures that enforce<br>relational geometry at initialization.</p> |
| title | Pay Attention Later: The Semantic Alignment Tax Requires Decoupling Representation from Reasoning |
| topic | Representation Learning Transformer Optimization Inductive Bias Semantic Geometry Initialization Schemes Low-Resource NLP Semantic Alignment Tax Deep learning Transformers Neuro-AI |
| url | https://doi.org/10.5281/zenodo.17667170 |