Beyond Parallelism: Synergistic Computational Graph Effects in Multi-Head Attention
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Borde, Haitz Sáez de Ocáriz |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
k-Maximum Inner Product Attention for Graph Transformers and the Expressive Power of GraphGPS
von: De Schouwer, Jonas, et al.
Veröffentlicht: (2026)
von: De Schouwer, Jonas, et al.
Veröffentlicht: (2026)
Neural Snowflakes: Universal Latent Graph Inference via Trainable Latent Geometries
von: Borde, Haitz Sáez de Ocáriz, et al.
Veröffentlicht: (2023)
von: Borde, Haitz Sáez de Ocáriz, et al.
Veröffentlicht: (2023)
Mathematical Foundations of Geometric Deep Learning
von: Borde, Haitz Sáez de Ocáriz, et al.
Veröffentlicht: (2025)
von: Borde, Haitz Sáez de Ocáriz, et al.
Veröffentlicht: (2025)
Fine-Tuning Next-Scale Visual Autoregressive Models with Group Relative Policy Optimization
von: Gallici, Matteo, et al.
Veröffentlicht: (2025)
von: Gallici, Matteo, et al.
Veröffentlicht: (2025)
Scalable Message Passing Neural Networks: No Need for Attention in Large Graph Representation Learning
von: Borde, Haitz Sáez de Ocáriz, et al.
Veröffentlicht: (2024)
von: Borde, Haitz Sáez de Ocáriz, et al.
Veröffentlicht: (2024)
Closed-Form Diffusion Models
von: Scarvelis, Christopher, et al.
Veröffentlicht: (2023)
von: Scarvelis, Christopher, et al.
Veröffentlicht: (2023)
Bridging Graph and State-Space Modeling for Intensive Care Unit Length of Stay Prediction
von: Zi, Shuqi, et al.
Veröffentlicht: (2025)
von: Zi, Shuqi, et al.
Veröffentlicht: (2025)
LoRA Fine-Tuning Without GPUs: A CPU-Efficient Meta-Generation Framework for LLMs
von: Arabpour, Reza, et al.
Veröffentlicht: (2025)
von: Arabpour, Reza, et al.
Veröffentlicht: (2025)
Towards Quantifying Long-Range Interactions in Graph Machine Learning: a Large Graph Dataset and a Measurement
von: Liang, Huidong, et al.
Veröffentlicht: (2025)
von: Liang, Huidong, et al.
Veröffentlicht: (2025)
Metric Learning for Clifford Group Equivariant Neural Networks
von: Ali, Riccardo, et al.
Veröffentlicht: (2024)
von: Ali, Riccardo, et al.
Veröffentlicht: (2024)
Approximation Rates and VC-Dimension Bounds for (P)ReLU MLP Mixture of Experts
von: Kratsios, Anastasis, et al.
Veröffentlicht: (2024)
von: Kratsios, Anastasis, et al.
Veröffentlicht: (2024)
Sharp Generalization Bounds for Foundation Models with Asymmetric Randomized Low-Rank Adapters
von: Kratsios, Anastasis, et al.
Veröffentlicht: (2025)
von: Kratsios, Anastasis, et al.
Veröffentlicht: (2025)
Classification Fields: Arbitrarily Fine Recursive Hierarchical Clustering From Few Examples
von: Li, Yicen, et al.
Veröffentlicht: (2026)
von: Li, Yicen, et al.
Veröffentlicht: (2026)
Neural Spacetimes for DAG Representation Learning
von: Borde, Haitz Sáez de Ocáriz, et al.
Veröffentlicht: (2024)
von: Borde, Haitz Sáez de Ocáriz, et al.
Veröffentlicht: (2024)
Keep It Light! Simplifying Image Clustering Via Text-Free Adapters
von: Li, Yicen, et al.
Veröffentlicht: (2025)
von: Li, Yicen, et al.
Veröffentlicht: (2025)
The Stepwise Informativeness Assumption: Why are Entropy Dynamics and Reasoning Correlated in LLMs?
von: Català, Mar Gonzàlez I, et al.
Veröffentlicht: (2026)
von: Català, Mar Gonzàlez I, et al.
Veröffentlicht: (2026)
Mitigating Model Drift in Developing Economies Using Synthetic Data and Outliers
von: Varshavskiy, Ilyas, et al.
Veröffentlicht: (2025)
von: Varshavskiy, Ilyas, et al.
Veröffentlicht: (2025)
Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M)
von: Merchant, Nicholas, et al.
Veröffentlicht: (2025)
von: Merchant, Nicholas, et al.
Veröffentlicht: (2025)
Every Feedforward Neural Network Definable in an o-Minimal Structure Has Finite Sample Complexity
von: Kratsios, Anastasis, et al.
Veröffentlicht: (2026)
von: Kratsios, Anastasis, et al.
Veröffentlicht: (2026)
Assessing the Geographic Generalization and Physical Consistency of Generative Models for Climate Downscaling
von: Saccardi, Carlo, et al.
Veröffentlicht: (2025)
von: Saccardi, Carlo, et al.
Veröffentlicht: (2025)
Score Distillation via Reparametrized DDIM
von: Lukoianov, Artem, et al.
Veröffentlicht: (2024)
von: Lukoianov, Artem, et al.
Veröffentlicht: (2024)
Asymmetry in Low-Rank Adapters of Foundation Models
von: Zhu, Jiacheng, et al.
Veröffentlicht: (2024)
von: Zhu, Jiacheng, et al.
Veröffentlicht: (2024)
Quantum Graph Attention Network: A Novel Quantum Multi-Head Attention Mechanism for Graph Learning
von: Ning, An, et al.
Veröffentlicht: (2025)
von: Ning, An, et al.
Veröffentlicht: (2025)
Multi-Head LatentMoE and Head Parallel: Communication-Efficient and Deterministic MoE Parallelism
von: Cui, Chenwei, et al.
Veröffentlicht: (2026)
von: Cui, Chenwei, et al.
Veröffentlicht: (2026)
Adaptive Head Budgeting for Efficient Multi-Head Attention
von: Faye, Bilal, et al.
Veröffentlicht: (2026)
von: Faye, Bilal, et al.
Veröffentlicht: (2026)
Learning 3D Hypersonic Flow with Physics-Enhanced Neural Fields: A Case Study on the Orion Reentry Capsule
von: Borde, Haitz Sáez de Ocáriz, et al.
Veröffentlicht: (2026)
von: Borde, Haitz Sáez de Ocáriz, et al.
Veröffentlicht: (2026)
DreamUp3D: Object-Centric Generative Models for Single-View 3D Scene Understanding and Real-to-Sim Transfer
von: Wu, Yizhe, et al.
Veröffentlicht: (2024)
von: Wu, Yizhe, et al.
Veröffentlicht: (2024)
Multi-Head Low-Rank Attention
von: Liu, Songtao, et al.
Veröffentlicht: (2026)
von: Liu, Songtao, et al.
Veröffentlicht: (2026)
The Effect of Attention Head Count on Transformer Approximation
von: Yu, Penghao, et al.
Veröffentlicht: (2025)
von: Yu, Penghao, et al.
Veröffentlicht: (2025)
Memorization Capacity of Multi-Head Attention in Transformers
von: Mahdavi, Sadegh, et al.
Veröffentlicht: (2023)
von: Mahdavi, Sadegh, et al.
Veröffentlicht: (2023)
Soro: A Lightweight Foundation Model and Chatbot for Tajik
von: Liashkov, Stanislav, et al.
Veröffentlicht: (2026)
von: Liashkov, Stanislav, et al.
Veröffentlicht: (2026)
MoH: Multi-Head Attention as Mixture-of-Head Attention
von: Jin, Peng, et al.
Veröffentlicht: (2024)
von: Jin, Peng, et al.
Veröffentlicht: (2024)
Beyond Classical Attention: Quantum Attention for Scalable Computation
von: Guo, Xuyang, et al.
Veröffentlicht: (2023)
von: Guo, Xuyang, et al.
Veröffentlicht: (2023)
Interleaved Head Attention
von: Duvvuri, Sai Surya, et al.
Veröffentlicht: (2026)
von: Duvvuri, Sai Surya, et al.
Veröffentlicht: (2026)
Boosting House Price Estimations with Multi-Head Gated Attention
von: Sellam, Zakaria Abdellah, et al.
Veröffentlicht: (2024)
von: Sellam, Zakaria Abdellah, et al.
Veröffentlicht: (2024)
MSPT: Efficient Large-Scale Physical Modeling via Parallelized Multi-Scale Attention
von: Curvo, Pedro M. P., et al.
Veröffentlicht: (2025)
von: Curvo, Pedro M. P., et al.
Veröffentlicht: (2025)
Geometric Analysis of Token Selection in Multi-Head Attention
von: Mudarisov, Timur, et al.
Veröffentlicht: (2026)
von: Mudarisov, Timur, et al.
Veröffentlicht: (2026)
Improving Transformers with Dynamically Composable Multi-Head Attention
von: Xiao, Da, et al.
Veröffentlicht: (2024)
von: Xiao, Da, et al.
Veröffentlicht: (2024)
Superiority of Multi-Head Attention in In-Context Linear Regression
von: Cui, Yingqian, et al.
Veröffentlicht: (2024)
von: Cui, Yingqian, et al.
Veröffentlicht: (2024)
Sparse Query Attention (SQA): A Computationally Efficient Attention Mechanism with Query Heads Reduction
von: Filipek, Adam
Veröffentlicht: (2025)
von: Filipek, Adam
Veröffentlicht: (2025)
Ähnliche Einträge
-
k-Maximum Inner Product Attention for Graph Transformers and the Expressive Power of GraphGPS
von: De Schouwer, Jonas, et al.
Veröffentlicht: (2026) -
Neural Snowflakes: Universal Latent Graph Inference via Trainable Latent Geometries
von: Borde, Haitz Sáez de Ocáriz, et al.
Veröffentlicht: (2023) -
Mathematical Foundations of Geometric Deep Learning
von: Borde, Haitz Sáez de Ocáriz, et al.
Veröffentlicht: (2025) -
Fine-Tuning Next-Scale Visual Autoregressive Models with Group Relative Policy Optimization
von: Gallici, Matteo, et al.
Veröffentlicht: (2025) -
Scalable Message Passing Neural Networks: No Need for Attention in Large Graph Representation Learning
von: Borde, Haitz Sáez de Ocáriz, et al.
Veröffentlicht: (2024)