Rank, Head-Channel Non-Identifiability, and Symmetry Breaking: A Precise Analysis of Representational Collapse in Transformers
Fuente:
arXiv
Saved in:
| Main Author: | Cirrincione, Giansalvo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the Geometry of Positional Encodings in Transformers
by: Cirrincione, Giansalvo
Published: (2026)
by: Cirrincione, Giansalvo
Published: (2026)
Hierarchical Kernel Transformer: Multi-Scale Attention with an Information-Theoretic Approximation Analysis
by: Cirrincione, Giansalvo
Published: (2026)
by: Cirrincione, Giansalvo
Published: (2026)
DDCL-INCRT: A Self-Organising Transformer with Hierarchical Prototype Structure (Theoretical Foundations)
by: Cirrincione, Giansalvo
Published: (2026)
by: Cirrincione, Giansalvo
Published: (2026)
DDCL: Deep Dual Competitive Learning: A Differentiable End-to-End Framework for Unsupervised Prototype-Based Representation Learning
by: Cirrincione, Giansalvo
Published: (2026)
by: Cirrincione, Giansalvo
Published: (2026)
INCRT: An Incremental Transformer That Determines Its Own Architecture
by: Cirrincione, Giansalvo
Published: (2026)
by: Cirrincione, Giansalvo
Published: (2026)
Collapse-Free Prototype Readout Layer for Transformer Encoders
by: Cirrincione, Giansalvo, et al.
Published: (2026)
by: Cirrincione, Giansalvo, et al.
Published: (2026)
Learned Lyapunov Shielding for Adaptive Control
by: Cirrincione, Giansalvo, et al.
Published: (2026)
by: Cirrincione, Giansalvo, et al.
Published: (2026)
Temporal Attention for Adaptive Control of Euler-Lagrange Systems with Unobservable Memory
by: Cirrincione, Giansalvo, et al.
Published: (2026)
by: Cirrincione, Giansalvo, et al.
Published: (2026)
Breaking Symmetry When Training Transformers
by: Zuo, Chunsheng, et al.
Published: (2024)
by: Zuo, Chunsheng, et al.
Published: (2024)
The geometry of BERT
by: Bonino, Matteo, et al.
Published: (2025)
by: Bonino, Matteo, et al.
Published: (2025)
Learning Adapter Rank via Symmetry Breaking
by: Doyle, Cooper, et al.
Published: (2025)
by: Doyle, Cooper, et al.
Published: (2025)
Seq-VCR: Preventing Collapse in Intermediate Transformer Representations for Enhanced Reasoning
by: Arefin, Md Rifat, et al.
Published: (2024)
by: Arefin, Md Rifat, et al.
Published: (2024)
Collapse of Irrelevant Representations (CIR) Ensures Robust and Non-Disruptive LLM Unlearning
by: Sondej, Filip, et al.
Published: (2025)
by: Sondej, Filip, et al.
Published: (2025)
SPARKLING: Balancing Signal Preservation and Symmetry Breaking for Width-Progressive Learning
by: Yu, Qifan, et al.
Published: (2026)
by: Yu, Qifan, et al.
Published: (2026)
Representation Collapse in Machine Translation Through the Lens of Angular Dispersion
by: Tokarchuk, Evgeniia, et al.
Published: (2026)
by: Tokarchuk, Evgeniia, et al.
Published: (2026)
Dynamic Rank Reinforcement Learning for Adaptive Low-Rank Multi-Head Self Attention in Large Language Models
by: Erden, Caner
Published: (2025)
by: Erden, Caner
Published: (2025)
SlimSpec: Low-Rank Draft LM-Head for Accelerated Speculative Decoding
by: Plaksin, Anton, et al.
Published: (2026)
by: Plaksin, Anton, et al.
Published: (2026)
Dimensional Collapse in Transformer Attention Outputs: A Challenge for Sparse Dictionary Learning
by: Wang, Junxuan, et al.
Published: (2025)
by: Wang, Junxuan, et al.
Published: (2025)
Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse
by: Fu, Zizhuo, et al.
Published: (2026)
by: Fu, Zizhuo, et al.
Published: (2026)
A Comparative Analysis of Contextual Representation Flow in State-Space and Transformer Architectures
by: Hoang, Nhat M., et al.
Published: (2025)
by: Hoang, Nhat M., et al.
Published: (2025)
Improving Transformers with Dynamically Composable Multi-Head Attention
by: Xiao, Da, et al.
Published: (2024)
by: Xiao, Da, et al.
Published: (2024)
Advancing LLM Safe Alignment with Safety Representation Ranking
by: Du, Tianqi, et al.
Published: (2025)
by: Du, Tianqi, et al.
Published: (2025)
Mechanism and Emergence of Stacked Attention Heads in Multi-Layer Transformers
by: Musat, Tiberiu
Published: (2024)
by: Musat, Tiberiu
Published: (2024)
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals
by: Saxena, Utkarsh, et al.
Published: (2024)
by: Saxena, Utkarsh, et al.
Published: (2024)
Contextually Guided Transformers via Low-Rank Adaptation
by: Zhmoginov, Andrey, et al.
Published: (2025)
by: Zhmoginov, Andrey, et al.
Published: (2025)
Breaking Quadratic Barriers: A Non-Attention LLM for Ultra-Long Context Horizons
by: Kiruluta, Andrew, et al.
Published: (2025)
by: Kiruluta, Andrew, et al.
Published: (2025)
Calibration Collapse Under Sycophancy Fine-Tuning: How Reward Hacking Breaks Uncertainty Quantification in LLMs
by: Sahoo, Subramanyam
Published: (2026)
by: Sahoo, Subramanyam
Published: (2026)
Transformers with Selective Access to Early Representations
by: Gunasekaran, Skye, et al.
Published: (2026)
by: Gunasekaran, Skye, et al.
Published: (2026)
Probabilistic Topic Modelling with Transformer Representations
by: Reuter, Arik, et al.
Published: (2024)
by: Reuter, Arik, et al.
Published: (2024)
GaLore$+$: Boosting Low-Rank Adaptation for LLMs with Cross-Head Projection
by: Liao, Xutao, et al.
Published: (2024)
by: Liao, Xutao, et al.
Published: (2024)
Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data
by: Gerstgrasser, Matthias, et al.
Published: (2024)
by: Gerstgrasser, Matthias, et al.
Published: (2024)
SwitchHead: Accelerating Transformers with Mixture-of-Experts Attention
by: Csordás, Róbert, et al.
Published: (2023)
by: Csordás, Róbert, et al.
Published: (2023)
Breaking BERT: Gradient Attack on Twitter Sentiment Analysis for Targeted Misclassification
by: Subedi, Akil Raj, et al.
Published: (2025)
by: Subedi, Akil Raj, et al.
Published: (2025)
DHA: Learning Decoupled-Head Attention from Transformer Checkpoints via Adaptive Heads Fusion
by: Chen, Yilong, et al.
Published: (2024)
by: Chen, Yilong, et al.
Published: (2024)
The Anxiety of Influence: Bloom Filters in Transformer Attention Heads
by: Balogh, Peter
Published: (2026)
by: Balogh, Peter
Published: (2026)
You Do Not Fully Utilize Transformer's Representation Capacity
by: Gerasimov, Gleb, et al.
Published: (2025)
by: Gerasimov, Gleb, et al.
Published: (2025)
ComplexFormer: Disruptively Advancing Transformer Inference Ability via Head-Specific Complex Vector Attention
by: Shao, Jintian, et al.
Published: (2025)
by: Shao, Jintian, et al.
Published: (2025)
The Expressive Power of Low Precision Softmax Transformers with (Summarized) Chain-of-Thought
by: Brösamle, Moritz, et al.
Published: (2026)
by: Brösamle, Moritz, et al.
Published: (2026)
Breaking the Attention Bottleneck
by: Hilsenbek, Kalle
Published: (2024)
by: Hilsenbek, Kalle
Published: (2024)
PMET: Precise Model Editing in a Transformer
by: Li, Xiaopeng, et al.
Published: (2023)
by: Li, Xiaopeng, et al.
Published: (2023)
Similar Items
-
On the Geometry of Positional Encodings in Transformers
by: Cirrincione, Giansalvo
Published: (2026) -
Hierarchical Kernel Transformer: Multi-Scale Attention with an Information-Theoretic Approximation Analysis
by: Cirrincione, Giansalvo
Published: (2026) -
DDCL-INCRT: A Self-Organising Transformer with Hierarchical Prototype Structure (Theoretical Foundations)
by: Cirrincione, Giansalvo
Published: (2026) -
DDCL: Deep Dual Competitive Learning: A Differentiable End-to-End Framework for Unsupervised Prototype-Based Representation Learning
by: Cirrincione, Giansalvo
Published: (2026) -
INCRT: An Incremental Transformer That Determines Its Own Architecture
by: Cirrincione, Giansalvo
Published: (2026)