Tucker Attention: A generalization of approximate attention mechanisms
Fuente:
arXiv
Guardado en:
| Autores principales: | Klein, Timon, Kusch, Jonas, Sager, Sebastian, Schnake, Stefan, Schotthöfer, Steffen |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
A geometric framework for momentum-based optimizers for low-rank training
por: Schotthöfer, Steffen, et al.
Publicado: (2025)
por: Schotthöfer, Steffen, et al.
Publicado: (2025)
Mitigating Subject Dependency in EEG Decoding with Subject-Specific Low-Rank Adapters
por: Klein, Timon, et al.
Publicado: (2025)
por: Klein, Timon, et al.
Publicado: (2025)
GeoLoRA: Geometric integration for parameter efficient fine-tuning
por: Schotthöfer, Steffen, et al.
Publicado: (2024)
por: Schotthöfer, Steffen, et al.
Publicado: (2024)
Geometry-aware training of factorized layers in tensor Tucker format
por: Zangrando, Emanuele, et al.
Publicado: (2023)
por: Zangrando, Emanuele, et al.
Publicado: (2023)
An Augmented Backward-Corrected Projector Splitting Integrator for Dynamical Low-Rank Training
por: Kusch, Jonas, et al.
Publicado: (2025)
por: Kusch, Jonas, et al.
Publicado: (2025)
Dynamical Low-Rank Compression of Neural Networks with Robustness under Adversarial Attacks
por: Schotthöfer, Steffen, et al.
Publicado: (2025)
por: Schotthöfer, Steffen, et al.
Publicado: (2025)
Federated Dynamical Low-Rank Training with Global Loss Convergence Guarantees
por: Schotthöfer, Steffen, et al.
Publicado: (2024)
por: Schotthöfer, Steffen, et al.
Publicado: (2024)
Compressing Vision Transformers in Geospatial Transfer Learning with Manifold-Constrained Optimization
por: Snyder, Thomas, et al.
Publicado: (2026)
por: Snyder, Thomas, et al.
Publicado: (2026)
HADL Framework for Noise Resilient Long-Term Time Series Forecasting
por: Dey, Aditya, et al.
Publicado: (2025)
por: Dey, Aditya, et al.
Publicado: (2025)
Poly-attention: a general scheme for higher-order self-attention
por: Chakrabarti, Sayak, et al.
Publicado: (2026)
por: Chakrabarti, Sayak, et al.
Publicado: (2026)
LORA-CRAFT: Cross-layer Rank Adaptation via Frozen Tucker Decomposition of Pre-trained Attention Weights
por: Dewage, Kasun, et al.
Publicado: (2026)
por: Dewage, Kasun, et al.
Publicado: (2026)
Towards Symbolic XAI -- Explanation Through Human Understandable Logical Relationships Between Features
por: Schnake, Thomas, et al.
Publicado: (2024)
por: Schnake, Thomas, et al.
Publicado: (2024)
Sink vs. diagonal patterns as mechanisms for attention switch and oversmoothing prevention
por: Súkeník, Peter, et al.
Publicado: (2026)
por: Súkeník, Peter, et al.
Publicado: (2026)
Physics-informed GNN for medium-high voltage AC power flow with edge-aware attention and line search correction operator
por: Kim, Changhun, et al.
Publicado: (2025)
por: Kim, Changhun, et al.
Publicado: (2025)
Attention mechanisms in neural networks
por: Hays, Hasi
Publicado: (2026)
por: Hays, Hasi
Publicado: (2026)
Uncovering the Structure of Explanation Quality with Spectral Analysis
por: Maeß, Johannes, et al.
Publicado: (2025)
por: Maeß, Johannes, et al.
Publicado: (2025)
Mode-Aware Non-Linear Tucker Autoencoder for Tensor-based Unsupervised Learning
por: Zheng, Junjing, et al.
Publicado: (2025)
por: Zheng, Junjing, et al.
Publicado: (2025)
Expanding Expressivity in Transformer Models with MöbiusAttention
por: Halacheva, Anna-Maria, et al.
Publicado: (2024)
por: Halacheva, Anna-Maria, et al.
Publicado: (2024)
Reorganizing attention-space geometry with expressive attention
por: Gros, Claudius
Publicado: (2024)
por: Gros, Claudius
Publicado: (2024)
A foundation model with multi-variate parallel attention to generate neuronal activity
por: Carzaniga, Francesco, et al.
Publicado: (2025)
por: Carzaniga, Francesco, et al.
Publicado: (2025)
Worst-case low-rank approximations
por: Fries, Anya, et al.
Publicado: (2026)
por: Fries, Anya, et al.
Publicado: (2026)
Efficiently Dispatching Flash Attention For Partially Filled Attention Masks
por: Sharma, Agniv, et al.
Publicado: (2024)
por: Sharma, Agniv, et al.
Publicado: (2024)
MEGAN: Multi-Explanation Graph Attention Network
por: Teufel, Jonas, et al.
Publicado: (2022)
por: Teufel, Jonas, et al.
Publicado: (2022)
Supervised learning pays attention
por: Craig, Erin, et al.
Publicado: (2025)
por: Craig, Erin, et al.
Publicado: (2025)
EEG motor imagery decoding: A framework for comparative analysis with channel attention mechanisms
por: Wimpff, Martin, et al.
Publicado: (2023)
por: Wimpff, Martin, et al.
Publicado: (2023)
Enhancing short-term traffic prediction by integrating trends and fluctuations with attention mechanism
por: Das, Adway, et al.
Publicado: (2025)
por: Das, Adway, et al.
Publicado: (2025)
A Diffusion-Based Method for Learning the Multi-Outcome Distribution of Medical Treatments
por: Ma, Yuchen, et al.
Publicado: (2025)
por: Ma, Yuchen, et al.
Publicado: (2025)
k-Maximum Inner Product Attention for Graph Transformers and the Expressive Power of GraphGPS
por: De Schouwer, Jonas, et al.
Publicado: (2026)
por: De Schouwer, Jonas, et al.
Publicado: (2026)
An end-to-end attention-based approach for learning on graphs
por: Buterez, David, et al.
Publicado: (2024)
por: Buterez, David, et al.
Publicado: (2024)
Short window attention enables long-term memorization
por: Cabannes, Loïc, et al.
Publicado: (2025)
por: Cabannes, Loïc, et al.
Publicado: (2025)
SortBench: Benchmarking LLMs based on their ability to sort lists
por: Herbold, Steffen
Publicado: (2025)
por: Herbold, Steffen
Publicado: (2025)
AlphaIntegrator: Transformer Action Search for Symbolic Integration Proofs
por: Ünsal, Mert, et al.
Publicado: (2024)
por: Ünsal, Mert, et al.
Publicado: (2024)
Depth-Recurrent Attention Mixtures: Giving Latent Reasoning the Attention it Deserves
por: Knupp, Jonas, et al.
Publicado: (2026)
por: Knupp, Jonas, et al.
Publicado: (2026)
CNN Explainability with Multivector Tucker Saliency Maps for Self-Supervised Models
por: Bouayed, Aymene Mohammed, et al.
Publicado: (2024)
por: Bouayed, Aymene Mohammed, et al.
Publicado: (2024)
Cross-attentive Cohesive Subgraph Embedding to Mitigate Oversquashing in GNNs
por: Hossain, Tanvir, et al.
Publicado: (2026)
por: Hossain, Tanvir, et al.
Publicado: (2026)
Twin Transformer using Gated Dynamic Learnable Attention mechanism for Fault Detection and Diagnosis in the Tennessee Eastman Process
por: Labbaf-Khaniki, Mohammad Ali, et al.
Publicado: (2024)
por: Labbaf-Khaniki, Mohammad Ali, et al.
Publicado: (2024)
Plasticity Loss in Deep Reinforcement Learning: A Survey
por: Klein, Timo, et al.
Publicado: (2024)
por: Klein, Timo, et al.
Publicado: (2024)
GQA-μP: The maximal parameterization update for grouped query attention
por: Chickering, Kyle R., et al.
Publicado: (2026)
por: Chickering, Kyle R., et al.
Publicado: (2026)
GRC-Net: Gram Residual Co-attention Net for epilepsy prediction
por: You, Bihao, et al.
Publicado: (2025)
por: You, Bihao, et al.
Publicado: (2025)
Static and multivariate-temporal attentive fusion transformer for readmission risk prediction
por: Sun, Zhe, et al.
Publicado: (2024)
por: Sun, Zhe, et al.
Publicado: (2024)
Ejemplares similares
-
A geometric framework for momentum-based optimizers for low-rank training
por: Schotthöfer, Steffen, et al.
Publicado: (2025) -
Mitigating Subject Dependency in EEG Decoding with Subject-Specific Low-Rank Adapters
por: Klein, Timon, et al.
Publicado: (2025) -
GeoLoRA: Geometric integration for parameter efficient fine-tuning
por: Schotthöfer, Steffen, et al.
Publicado: (2024) -
Geometry-aware training of factorized layers in tensor Tucker format
por: Zangrando, Emanuele, et al.
Publicado: (2023) -
An Augmented Backward-Corrected Projector Splitting Integrator for Dynamical Low-Rank Training
por: Kusch, Jonas, et al.
Publicado: (2025)