Laplacian Heads Improve Transformers by Smoothing Token Representations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Yuchong, Papyan, Vardan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On the Importance of Gaussianizing Representations
von: Eftekhari, Daniel, et al.
Veröffentlicht: (2025)
von: Eftekhari, Daniel, et al.
Veröffentlicht: (2025)
OATS: Outlier-Aware Pruning Through Sparse and Low Rank Decomposition
von: Zhang, Stephen, et al.
Veröffentlicht: (2024)
von: Zhang, Stephen, et al.
Veröffentlicht: (2024)
Sparsest Models Elude Pruning: An Exposé of Pruning's Current Capabilities
von: Zhang, Stephen, et al.
Veröffentlicht: (2024)
von: Zhang, Stephen, et al.
Veröffentlicht: (2024)
Residual Alignment: Uncovering the Mechanisms of Residual Networks
von: Li, Jianing, et al.
Veröffentlicht: (2024)
von: Li, Jianing, et al.
Veröffentlicht: (2024)
Pushing Boundaries: Mixup's Influence on Neural Collapse
von: Fisher, Quinn, et al.
Veröffentlicht: (2024)
von: Fisher, Quinn, et al.
Veröffentlicht: (2024)
Attention Sinks: A 'Catch, Tag, Release' Mechanism for Embeddings
von: Zhang, Stephen, et al.
Veröffentlicht: (2025)
von: Zhang, Stephen, et al.
Veröffentlicht: (2025)
Transformer Block Coupling and its Correlation with Generalization in LLMs
von: Aubry, Murdock, et al.
Veröffentlicht: (2024)
von: Aubry, Murdock, et al.
Veröffentlicht: (2024)
Linguistic Collapse: Neural Collapse in (Large) Language Models
von: Wu, Robert, et al.
Veröffentlicht: (2024)
von: Wu, Robert, et al.
Veröffentlicht: (2024)
Out of the Ordinary: Spectrally Adapting Regression for Covariate Shift
von: Eyre, Benjamin, et al.
Veröffentlicht: (2023)
von: Eyre, Benjamin, et al.
Veröffentlicht: (2023)
Towards Improved Sentence Representations using Token Graphs
von: Mantri, Krishna Sri Ipsit, et al.
Veröffentlicht: (2026)
von: Mantri, Krishna Sri Ipsit, et al.
Veröffentlicht: (2026)
Laplacian Representations for Decision-Time Planning
von: Shehmar, Dikshant, et al.
Veröffentlicht: (2026)
von: Shehmar, Dikshant, et al.
Veröffentlicht: (2026)
Leveraging Contrastive Learning for Enhanced Node Representations in Tokenized Graph Transformers
von: Chen, Jinsong, et al.
Veröffentlicht: (2024)
von: Chen, Jinsong, et al.
Veröffentlicht: (2024)
Improving Transformers with Dynamically Composable Multi-Head Attention
von: Xiao, Da, et al.
Veröffentlicht: (2024)
von: Xiao, Da, et al.
Veröffentlicht: (2024)
Proper Laplacian Representation Learning
von: Gomez, Diego, et al.
Veröffentlicht: (2023)
von: Gomez, Diego, et al.
Veröffentlicht: (2023)
Towards Stable, Globally Expressive Graph Representations with Laplacian Eigenvectors
von: Zhou, Junru, et al.
Veröffentlicht: (2024)
von: Zhou, Junru, et al.
Veröffentlicht: (2024)
Impact of Connectivity on Laplacian Representations in Reinforcement Learning
von: Giorgi, Tommaso, et al.
Veröffentlicht: (2026)
von: Giorgi, Tommaso, et al.
Veröffentlicht: (2026)
TokenBlowUp: Resolving Representational Singularities in LLM Token Spaces via Monoidal Transformations
von: Zhao, Dongfang
Veröffentlicht: (2025)
von: Zhao, Dongfang
Veröffentlicht: (2025)
GANs as Gradient Flows that Converge
von: Huang, Yu-Jui, et al.
Veröffentlicht: (2022)
von: Huang, Yu-Jui, et al.
Veröffentlicht: (2022)
Improving Random Forests by Smoothing
von: Liu, Ziyi, et al.
Veröffentlicht: (2025)
von: Liu, Ziyi, et al.
Veröffentlicht: (2025)
Functional Autoencoder for Smoothing and Representation Learning
von: Wu, Sidi, et al.
Veröffentlicht: (2024)
von: Wu, Sidi, et al.
Veröffentlicht: (2024)
Improving Model Representation and Reducing KV Cache via Skip Connections with First Value Heads
von: Wu, Zhoutong, et al.
Veröffentlicht: (2025)
von: Wu, Zhoutong, et al.
Veröffentlicht: (2025)
Supra-Laplacian Encoding for Transformer on Dynamic Graphs
von: Karmim, Yannis, et al.
Veröffentlicht: (2024)
von: Karmim, Yannis, et al.
Veröffentlicht: (2024)
Differential Smoothing Mitigates Sharpening and Improves LLM Reasoning
von: Gai, Jingchu, et al.
Veröffentlicht: (2025)
von: Gai, Jingchu, et al.
Veröffentlicht: (2025)
Double Self-weighted Multi-view Clustering via Adaptive View Fusion
von: Fang, Xiang, et al.
Veröffentlicht: (2020)
von: Fang, Xiang, et al.
Veröffentlicht: (2020)
A Concept-Centric Approach to Multi-Modality Learning
von: Geng, Yuchong, et al.
Veröffentlicht: (2024)
von: Geng, Yuchong, et al.
Veröffentlicht: (2024)
Label Smoothing Improves Machine Unlearning
von: Di, Zonglin, et al.
Veröffentlicht: (2024)
von: Di, Zonglin, et al.
Veröffentlicht: (2024)
Speech Tokenizer is Key to Consistent Representation
von: Jung, Wonjin, et al.
Veröffentlicht: (2025)
von: Jung, Wonjin, et al.
Veröffentlicht: (2025)
Lost in Tokenization: Fundamental Trade-offs in Graph Tokenization for Transformers
von: Bechler-Speicher, Maya, et al.
Veröffentlicht: (2026)
von: Bechler-Speicher, Maya, et al.
Veröffentlicht: (2026)
TokenFormer: Rethinking Transformer Scaling with Tokenized Model Parameters
von: Wang, Haiyang, et al.
Veröffentlicht: (2024)
von: Wang, Haiyang, et al.
Veröffentlicht: (2024)
Geometric Analysis of Token Selection in Multi-Head Attention
von: Mudarisov, Timur, et al.
Veröffentlicht: (2026)
von: Mudarisov, Timur, et al.
Veröffentlicht: (2026)
Improving Time Series Classification with Representation Soft Label Smoothing
von: Ma, Hengyi, et al.
Veröffentlicht: (2024)
von: Ma, Hengyi, et al.
Veröffentlicht: (2024)
Exploring Pseudo-Token Approaches in Transformer Neural Processes
von: Lara-Rangel, Jose, et al.
Veröffentlicht: (2025)
von: Lara-Rangel, Jose, et al.
Veröffentlicht: (2025)
Active-Dormant Attention Heads: Mechanistically Demystifying Extreme-Token Phenomena in LLMs
von: Guo, Tianyu, et al.
Veröffentlicht: (2024)
von: Guo, Tianyu, et al.
Veröffentlicht: (2024)
Complex-Valued Unitary Representations as Classification Heads for Improved Uncertainty Quantification in Deep Neural Networks
von: Jafari, Akbar Anbar, et al.
Veröffentlicht: (2026)
von: Jafari, Akbar Anbar, et al.
Veröffentlicht: (2026)
Laplacian Convolutional Representation for Traffic Time Series Imputation
von: Chen, Xinyu, et al.
Veröffentlicht: (2022)
von: Chen, Xinyu, et al.
Veröffentlicht: (2022)
Topology-guided Hypergraph Transformer Network: Unveiling Structural Insights for Improved Representation
von: Saifuddin, Khaled Mohammed, et al.
Veröffentlicht: (2023)
von: Saifuddin, Khaled Mohammed, et al.
Veröffentlicht: (2023)
Probabilistic Smoothing with Ratio-Monotone Transforms for Global Optimization
von: Jang, Kukyoung, et al.
Veröffentlicht: (2026)
von: Jang, Kukyoung, et al.
Veröffentlicht: (2026)
Memorization Capacity of Multi-Head Attention in Transformers
von: Mahdavi, Sadegh, et al.
Veröffentlicht: (2023)
von: Mahdavi, Sadegh, et al.
Veröffentlicht: (2023)
The Effect of Attention Head Count on Transformer Approximation
von: Yu, Penghao, et al.
Veröffentlicht: (2025)
von: Yu, Penghao, et al.
Veröffentlicht: (2025)
A Theoretical Framework for Prompt Engineering: Approximating Smooth Functions with Transformer Prompts
von: Nakada, Ryumei, et al.
Veröffentlicht: (2025)
von: Nakada, Ryumei, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
On the Importance of Gaussianizing Representations
von: Eftekhari, Daniel, et al.
Veröffentlicht: (2025) -
OATS: Outlier-Aware Pruning Through Sparse and Low Rank Decomposition
von: Zhang, Stephen, et al.
Veröffentlicht: (2024) -
Sparsest Models Elude Pruning: An Exposé of Pruning's Current Capabilities
von: Zhang, Stephen, et al.
Veröffentlicht: (2024) -
Residual Alignment: Uncovering the Mechanisms of Residual Networks
von: Li, Jianing, et al.
Veröffentlicht: (2024) -
Pushing Boundaries: Mixup's Influence on Neural Collapse
von: Fisher, Quinn, et al.
Veröffentlicht: (2024)