Learning Advanced Self-Attention for Linear Transformers in the Singular Value Domain
Fuente:
arXiv
Saved in:
| Main Authors: | Wi, Hyowon, Choi, Jeongwhan, Park, Noseong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
An Attentive Inductive Bias for Sequential Recommendation beyond the Self-Attention
by: Shin, Yehjin, et al.
Published: (2023)
by: Shin, Yehjin, et al.
Published: (2023)
Graph Convolutions Enrich the Self-Attention in Transformers!
by: Choi, Jeongwhan, et al.
Published: (2023)
by: Choi, Jeongwhan, et al.
Published: (2023)
PANDA: Expanded Width-Aware Message Passing Beyond Rewiring
by: Choi, Jeongwhan, et al.
Published: (2024)
by: Choi, Jeongwhan, et al.
Published: (2024)
SCONE: A Novel Stochastic Sampling to Generate Contrastive Views and Hard Negative Samples for Recommendation
by: Lee, Chaejeong, et al.
Published: (2024)
by: Lee, Chaejeong, et al.
Published: (2024)
RDGCL: Reaction-Diffusion Graph Contrastive Learning for Recommendation
by: Choi, Jeongwhan, et al.
Published: (2023)
by: Choi, Jeongwhan, et al.
Published: (2023)
Are Graph Transformers Necessary? Efficient Long-Range Message Passing with Fractal Nodes in MPNNs
by: Choi, Jeongwhan, et al.
Published: (2025)
by: Choi, Jeongwhan, et al.
Published: (2025)
TV-Rec: Time-Variant Convolutional Filter for Sequential Recommendation
by: Shin, Yehjin, et al.
Published: (2025)
by: Shin, Yehjin, et al.
Published: (2025)
Can TabPFN Compete with GNNs for Node Classification via Graph Tabularization?
by: Choi, Jeongwhan, et al.
Published: (2025)
by: Choi, Jeongwhan, et al.
Published: (2025)
PIORF: Physics-Informed Ollivier-Ricci Flow for Long-Range Interactions in Mesh Graph Neural Networks
by: Yu, Youn-Yeol, et al.
Published: (2025)
by: Yu, Youn-Yeol, et al.
Published: (2025)
Possibility for Proactive Anomaly Detection
by: Jeon, Jinsung, et al.
Published: (2025)
by: Jeon, Jinsung, et al.
Published: (2025)
SVD-AE: Simple Autoencoders for Collaborative Filtering
by: Hong, Seoyoung, et al.
Published: (2024)
by: Hong, Seoyoung, et al.
Published: (2024)
Continuous-time Autoencoders for Regular and Irregular Time Series Imputation
by: Wi, Hyowon, et al.
Published: (2023)
by: Wi, Hyowon, et al.
Published: (2023)
Bridging Dynamic Factor Models and Neural Controlled Differential Equations for Nowcasting GDP
by: Lim, Seonkyu, et al.
Published: (2024)
by: Lim, Seonkyu, et al.
Published: (2024)
How Many Experts Are Enough? Towards Optimal Semantic Specialization for Mixture-of-Experts
by: Park, Sumin, et al.
Published: (2025)
by: Park, Sumin, et al.
Published: (2025)
Learning Posterior Predictive Distributions for Node Classification from Synthetic Graph Priors
by: Choi, Jeongwhan, et al.
Published: (2026)
by: Choi, Jeongwhan, et al.
Published: (2026)
SPI-GAN: Denoising Diffusion GANs with Straight-Path Interpolations
by: Jeon, Jinsung, et al.
Published: (2022)
by: Jeon, Jinsung, et al.
Published: (2022)
HINTS: Extraction of Human Insights from Time-Series Without External Sources
by: Jhin, Sheo Yon, et al.
Published: (2025)
by: Jhin, Sheo Yon, et al.
Published: (2025)
Neural Functions for Learning Periodic Signal
by: Cho, Woojin, et al.
Published: (2025)
by: Cho, Woojin, et al.
Published: (2025)
MimiQ: Low-Bit Data-Free Quantization of Vision Transformers with Encouraging Inter-Head Attention Similarity
by: Choi, Kanghyun, et al.
Published: (2024)
by: Choi, Kanghyun, et al.
Published: (2024)
Learning Flexible Body Collision Dynamics with Hierarchical Contact Mesh Transformer
by: Yu, Youn-Yeol, et al.
Published: (2023)
by: Yu, Youn-Yeol, et al.
Published: (2023)
Graph Signal Processing Meets Mamba2: Adaptive Filter Bank via Delta Modulation
by: Shin, Yehjin, et al.
Published: (2026)
by: Shin, Yehjin, et al.
Published: (2026)
VSFormer: Value and Shape-Aware Transformer with Prior-Enhanced Self-Attention for Multivariate Time Series Classification
by: Xi, Wenjie, et al.
Published: (2024)
by: Xi, Wenjie, et al.
Published: (2024)
Towards Unified and Adaptive Cross-Domain Collaborative Filtering via Graph Signal Processing
by: Lee, Jeongeun, et al.
Published: (2024)
by: Lee, Jeongeun, et al.
Published: (2024)
Transolver is a Linear Transformer: Revisiting Physics-Attention through the Lens of Linear Attention
by: Hu, Wenjie, et al.
Published: (2025)
by: Hu, Wenjie, et al.
Published: (2025)
Key and Value Weights Are Probably All You Need: On the Necessity of the Query, Key, Value weight Triplet in Self-Attention Transformers
by: Karbevski, Marko, et al.
Published: (2025)
by: Karbevski, Marko, et al.
Published: (2025)
Attention-Aided MMSE for OFDM Channel Estimation: Learning Linear Filters with Attention
by: Ha, TaeJun, et al.
Published: (2025)
by: Ha, TaeJun, et al.
Published: (2025)
Cross-Architecture Transfer Learning for Linear-Cost Inference Transformers
by: Choi, Sehyun
Published: (2024)
by: Choi, Sehyun
Published: (2024)
ZeroS: Zero-Sum Linear Attention for Efficient Transformers
by: Lu, Jiecheng, et al.
Published: (2026)
by: Lu, Jiecheng, et al.
Published: (2026)
Holographic Transformers for Complex-Valued Signal Processing: Integrating Phase Interference into Self-Attention
by: Huang, Enhao, et al.
Published: (2025)
by: Huang, Enhao, et al.
Published: (2025)
Enhancing Linear Attention with Residual Learning
by: Lai, Xunhao, et al.
Published: (2025)
by: Lai, Xunhao, et al.
Published: (2025)
Self-Supervised Contrastive Learning for Long-term Forecasting
by: Park, Junwoo, et al.
Published: (2024)
by: Park, Junwoo, et al.
Published: (2024)
FANoise: Singular Value-Adaptive Noise Modulation for Robust Multimodal Representation Learning
by: Li, Jiaoyang, et al.
Published: (2025)
by: Li, Jiaoyang, et al.
Published: (2025)
PiSSA: Principal Singular Values and Singular Vectors Adaptation of Large Language Models
by: Meng, Fanxu, et al.
Published: (2024)
by: Meng, Fanxu, et al.
Published: (2024)
Attention-based Iterative Decomposition for Tensor Product Representation
by: Park, Taewon, et al.
Published: (2024)
by: Park, Taewon, et al.
Published: (2024)
Value-State Gated Attention for Mitigating Extreme-Token Phenomena in Transformers
by: Bu, Rui, et al.
Published: (2025)
by: Bu, Rui, et al.
Published: (2025)
Singular Vectors of Attention Heads Align with Features
by: Franco, Gabriel, et al.
Published: (2026)
by: Franco, Gabriel, et al.
Published: (2026)
A Generalized Singular Value Theory for Neural Networks
by: Brown, Brian Charles, et al.
Published: (2026)
by: Brown, Brian Charles, et al.
Published: (2026)
Are Self-Attentions Effective for Time Series Forecasting?
by: Kim, Dongbin, et al.
Published: (2024)
by: Kim, Dongbin, et al.
Published: (2024)
MaD-Scientist: AI-based Scientist solving Convection-Diffusion-Reaction Equations Using Massive PINN-Based Prior Data
by: Kang, Mingu, et al.
Published: (2024)
by: Kang, Mingu, et al.
Published: (2024)
Eigen-Value: Efficient Domain-Robust Data Valuation via Eigenvalue-Based Approach
by: Choi, Youngjun, et al.
Published: (2025)
by: Choi, Youngjun, et al.
Published: (2025)
Similar Items
-
An Attentive Inductive Bias for Sequential Recommendation beyond the Self-Attention
by: Shin, Yehjin, et al.
Published: (2023) -
Graph Convolutions Enrich the Self-Attention in Transformers!
by: Choi, Jeongwhan, et al.
Published: (2023) -
PANDA: Expanded Width-Aware Message Passing Beyond Rewiring
by: Choi, Jeongwhan, et al.
Published: (2024) -
SCONE: A Novel Stochastic Sampling to Generate Contrastive Views and Hard Negative Samples for Recommendation
by: Lee, Chaejeong, et al.
Published: (2024) -
RDGCL: Reaction-Diffusion Graph Contrastive Learning for Recommendation
by: Choi, Jeongwhan, et al.
Published: (2023)