Mind the Gap: a Spectral Analysis of Rank Collapse and Signal Propagation in Attention Layers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Saada, Thiziri Nait, Naderi, Alireza, Tanner, Jared |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Beyond IID weights: sparse and low-rank deep Neural Networks are also Gaussian Processes
von: Nait-Saada, Thiziri, et al.
Veröffentlicht: (2023)
von: Nait-Saada, Thiziri, et al.
Veröffentlicht: (2023)
A simple proof of almost sure convergence for the largest singular value of a product of Gaussian matrices
von: Saada, Thiziri Nait, et al.
Veröffentlicht: (2024)
von: Saada, Thiziri Nait, et al.
Veröffentlicht: (2024)
The Data-Quality Illusion: Rethinking Classifier-Based Quality Filtering for LLM Pretraining
von: Saada, Thiziri Nait, et al.
Veröffentlicht: (2025)
von: Saada, Thiziri Nait, et al.
Veröffentlicht: (2025)
Theory of Minimal Weight Perturbations in Deep Networks and its Applications for Low-Rank Activated Backdoor Attacks
von: Evans, Bethan, et al.
Veröffentlicht: (2026)
von: Evans, Bethan, et al.
Veröffentlicht: (2026)
GraphEval: A Knowledge-Graph Based LLM Hallucination Evaluation Framework
von: Sansford, Hannah, et al.
Veröffentlicht: (2024)
von: Sansford, Hannah, et al.
Veröffentlicht: (2024)
On the Benefits of Rank in Attention Layers
von: Amsel, Noah, et al.
Veröffentlicht: (2024)
von: Amsel, Noah, et al.
Veröffentlicht: (2024)
LOTFormer: Doubly-Stochastic Linear Attention via Low-Rank Optimal Transport
von: Shahbazi, Ashkan, et al.
Veröffentlicht: (2025)
von: Shahbazi, Ashkan, et al.
Veröffentlicht: (2025)
How Controlling the Variance can Improve Training Stability of Sparsely Activated DNNs and CNNs
von: Dent, Emily, et al.
Veröffentlicht: (2026)
von: Dent, Emily, et al.
Veröffentlicht: (2026)
CARMIL: Context-Aware Regularization on Multiple Instance Learning models for Whole Slide Images
von: Saada, Thiziri Nait, et al.
Veröffentlicht: (2024)
von: Saada, Thiziri Nait, et al.
Veröffentlicht: (2024)
Mind the Links: Cross-Layer Attention for Link Prediction in Multiplex Networks
von: Sharma, Devesh, et al.
Veröffentlicht: (2025)
von: Sharma, Devesh, et al.
Veröffentlicht: (2025)
Safe LLM-Controlled Robots with Formal Guarantees via Reachability Analysis
von: Hafez, Ahmad, et al.
Veröffentlicht: (2025)
von: Hafez, Ahmad, et al.
Veröffentlicht: (2025)
Quantifying Error Propagation and Model Collapse in Diffusion Models
von: Khelifa, Nail B., et al.
Veröffentlicht: (2026)
von: Khelifa, Nail B., et al.
Veröffentlicht: (2026)
Reinforcement Learning from LLM Feedback to Counteract Goal Misgeneralization
von: Barj, Houda Nait El, et al.
Veröffentlicht: (2024)
von: Barj, Houda Nait El, et al.
Veröffentlicht: (2024)
Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse
von: Fu, Zizhuo, et al.
Veröffentlicht: (2026)
von: Fu, Zizhuo, et al.
Veröffentlicht: (2026)
From Condensation to Rank Collapse: A Two-Stage Analysis of Transformer Training Dynamics
von: Chen, Zheng-An, et al.
Veröffentlicht: (2025)
von: Chen, Zheng-An, et al.
Veröffentlicht: (2025)
Layer Collapse in Diffusion Language Models
von: Conzelmann, Alexander, et al.
Veröffentlicht: (2026)
von: Conzelmann, Alexander, et al.
Veröffentlicht: (2026)
Cross-Modal Bayesian Low-Rank Adaptation for Uncertainty-Aware Multimodal Learning
von: Naderi, Habibeh, et al.
Veröffentlicht: (2026)
von: Naderi, Habibeh, et al.
Veröffentlicht: (2026)
Mind the Gap: Optimal and Equitable Encouragement Policies
von: Zhou, Angela
Veröffentlicht: (2023)
von: Zhou, Angela
Veröffentlicht: (2023)
The Persistence of Neural Collapse Despite Low-Rank Bias
von: Garrod, Connall, et al.
Veröffentlicht: (2024)
von: Garrod, Connall, et al.
Veröffentlicht: (2024)
Rank-Aware Spectral Bounds on Attention Logits for Stable Low-Precision Training
von: Emadi, Seyed Morteza
Veröffentlicht: (2026)
von: Emadi, Seyed Morteza
Veröffentlicht: (2026)
Mind the Gap: Removing the Discretization Gap in Differentiable Logic Gate Networks
von: Yousefi, Shakir, et al.
Veröffentlicht: (2025)
von: Yousefi, Shakir, et al.
Veröffentlicht: (2025)
Posterior Collapse as Automatic Spectral Pruning
von: Hirn, Johannes
Veröffentlicht: (2026)
von: Hirn, Johannes
Veröffentlicht: (2026)
Low-Rank Learning by Design: the Role of Network Architecture and Activation Linearity in Gradient Rank Collapse
von: Baker, Bradley T., et al.
Veröffentlicht: (2024)
von: Baker, Bradley T., et al.
Veröffentlicht: (2024)
Mind the Gap: Structure-Aware Consistency in Preference Learning
von: Mohri, Mehryar, et al.
Veröffentlicht: (2026)
von: Mohri, Mehryar, et al.
Veröffentlicht: (2026)
Mind the Gaps: Measuring Visual Artifacts in Dimensionality Reduction
von: Ros, Jaume, et al.
Veröffentlicht: (2025)
von: Ros, Jaume, et al.
Veröffentlicht: (2025)
Lambda-Skip Connections: the architectural component that prevents Rank Collapse
von: Joseph, Federico Arangath, et al.
Veröffentlicht: (2024)
von: Joseph, Federico Arangath, et al.
Veröffentlicht: (2024)
Preventing Representational Rank Collapse in MPNNs by Splitting the Computational Graph
von: Roth, Andreas, et al.
Veröffentlicht: (2024)
von: Roth, Andreas, et al.
Veröffentlicht: (2024)
Gaussian Equivalence for Self-Attention: Asymptotic Spectral Analysis of Attention Matrix
von: Hayase, Tomohiro, et al.
Veröffentlicht: (2025)
von: Hayase, Tomohiro, et al.
Veröffentlicht: (2025)
Model Merging by Output-Space Projection
von: Evans, Bethan, et al.
Veröffentlicht: (2026)
von: Evans, Bethan, et al.
Veröffentlicht: (2026)
When Attention Collapses: How Degenerate Layers in LLMs Enable Smaller, Stronger Models
von: Sanyal, Sunny, et al.
Veröffentlicht: (2024)
von: Sanyal, Sunny, et al.
Veröffentlicht: (2024)
DDP-SA: Scalable Privacy-Preserving Federated Learning via Distributed Differential Privacy and Secure Aggregation
von: Wei, Wenjing, et al.
Veröffentlicht: (2026)
von: Wei, Wenjing, et al.
Veröffentlicht: (2026)
How Label Imbalance Shapes Geometry: A General Spectral Analysis of Multi-Label Neural Collapse
von: Ma, Xiaoxuan, et al.
Veröffentlicht: (2026)
von: Ma, Xiaoxuan, et al.
Veröffentlicht: (2026)
Mind the Gap: Measuring Generalization Performance Across Multiple Objectives
von: Feurer, Matthias, et al.
Veröffentlicht: (2022)
von: Feurer, Matthias, et al.
Veröffentlicht: (2022)
The Inlet Rank Collapse in Implicit Neural Representations: Diagnosis and Unified Remedy
von: Zheng, Jianqiao, et al.
Veröffentlicht: (2026)
von: Zheng, Jianqiao, et al.
Veröffentlicht: (2026)
LOSS-GAT: Label Propagation and One-Class Semi-Supervised Graph Attention Network for Fake News Detection
von: Lakzaei, Batool, et al.
Veröffentlicht: (2024)
von: Lakzaei, Batool, et al.
Veröffentlicht: (2024)
Mind The Gap: Quantifying Mechanistic Gaps in Algorithmic Reasoning via Neural Compilation
von: Saldyt, Lucas, et al.
Veröffentlicht: (2025)
von: Saldyt, Lucas, et al.
Veröffentlicht: (2025)
LaCoOT: Layer Collapse through Optimal Transport
von: Quétu, Victor, et al.
Veröffentlicht: (2024)
von: Quétu, Victor, et al.
Veröffentlicht: (2024)
Subcritical Signal Propagation at Initialization in Normalization-Free Transformers
von: Alekseev, Sergey
Veröffentlicht: (2026)
von: Alekseev, Sergey
Veröffentlicht: (2026)
Layer Collapse Can be Induced by Unstructured Pruning
von: Liao, Zhu, et al.
Veröffentlicht: (2024)
von: Liao, Zhu, et al.
Veröffentlicht: (2024)
LayerCollapse: Adaptive compression of neural networks
von: Shabgahi, Soheil Zibakhsh, et al.
Veröffentlicht: (2023)
von: Shabgahi, Soheil Zibakhsh, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Beyond IID weights: sparse and low-rank deep Neural Networks are also Gaussian Processes
von: Nait-Saada, Thiziri, et al.
Veröffentlicht: (2023) -
A simple proof of almost sure convergence for the largest singular value of a product of Gaussian matrices
von: Saada, Thiziri Nait, et al.
Veröffentlicht: (2024) -
The Data-Quality Illusion: Rethinking Classifier-Based Quality Filtering for LLM Pretraining
von: Saada, Thiziri Nait, et al.
Veröffentlicht: (2025) -
Theory of Minimal Weight Perturbations in Deep Networks and its Applications for Low-Rank Activated Backdoor Attacks
von: Evans, Bethan, et al.
Veröffentlicht: (2026) -
GraphEval: A Knowledge-Graph Based LLM Hallucination Evaluation Framework
von: Sansford, Hannah, et al.
Veröffentlicht: (2024)