Self-Attention at Constant Cost per Token via Symmetry-Aware Taylor Approximation
Fuente:
arXiv
Salvato in:
| Autori principali: | Heinsen, Franz A., Kozachkov, Leo |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SATER: A Self-Aware and Token-Efficient Approach to Routing and Cascading
di: Shen, Yuanzhe, et al.
Pubblicazione: (2025)
di: Shen, Yuanzhe, et al.
Pubblicazione: (2025)
ATTENTION2D: Communication Efficient Distributed Self-Attention Mechanism
di: Elango, Venmugil
Pubblicazione: (2025)
di: Elango, Venmugil
Pubblicazione: (2025)
Cost-Efficient LLM Training with Lifetime-Aware Tensor Offloading via GPUDirect Storage
di: Yuan, Ziqi, et al.
Pubblicazione: (2025)
di: Yuan, Ziqi, et al.
Pubblicazione: (2025)
Cost-Efficient Multimodal LLM Inference via Cross-Tier GPU Heterogeneity
di: Yu, Donglin
Pubblicazione: (2026)
di: Yu, Donglin
Pubblicazione: (2026)
Token-Budget-Aware Pool Routing for Cost-Efficient LLM Inference
di: Chen, Huamin, et al.
Pubblicazione: (2026)
di: Chen, Huamin, et al.
Pubblicazione: (2026)
Context Parallelism for Scalable Million-Token Inference
di: Yang, Amy, et al.
Pubblicazione: (2024)
di: Yang, Amy, et al.
Pubblicazione: (2024)
SimulCost: A Cost-Aware Benchmark and Toolkit for Automating Physics Simulations with LLMs
di: Cao, Yadi, et al.
Pubblicazione: (2026)
di: Cao, Yadi, et al.
Pubblicazione: (2026)
Context-Aware Inference via Performance Forecasting in Decentralized Learning Networks
di: Pfeffer, Joel, et al.
Pubblicazione: (2025)
di: Pfeffer, Joel, et al.
Pubblicazione: (2025)
MAC-Attention: a Match-Amend-Complete Scheme for Fast and Accurate Attention Computation
di: Yao, Jinghan, et al.
Pubblicazione: (2026)
di: Yao, Jinghan, et al.
Pubblicazione: (2026)
Inference Offloading for Cost-Sensitive Binary Classification at the Edge
di: Moothedath, Vishnu Narayanan, et al.
Pubblicazione: (2025)
di: Moothedath, Vishnu Narayanan, et al.
Pubblicazione: (2025)
A Fast and Flat Federated Learning Method via Weighted Momentum and Sharpness-Aware Minimization
di: Li, Tianle, et al.
Pubblicazione: (2025)
di: Li, Tianle, et al.
Pubblicazione: (2025)
Stochastic Sparse Attention for Memory-Bound Inference
di: Lee, Kyle, et al.
Pubblicazione: (2026)
di: Lee, Kyle, et al.
Pubblicazione: (2026)
Joint Model Assignment and Resource Allocation for Cost-Effective Mobile Generative Services
di: Gao, Shuangwei, et al.
Pubblicazione: (2024)
di: Gao, Shuangwei, et al.
Pubblicazione: (2024)
DeepHYDRA: Resource-Efficient Time-Series Anomaly Detection in Dynamically-Configured Systems
di: Stehle, Franz Kevin, et al.
Pubblicazione: (2024)
di: Stehle, Franz Kevin, et al.
Pubblicazione: (2024)
EPSILON: Adaptive Fault Mitigation in Approximate Deep Neural Network using Statistical Signatures
di: Khalil, Khurram, et al.
Pubblicazione: (2025)
di: Khalil, Khurram, et al.
Pubblicazione: (2025)
The Illusion of Power Capping in LLM Decode: A Phase-Aware Energy Characterisation Across Attention Architectures
di: Ma, Bole, et al.
Pubblicazione: (2026)
di: Ma, Bole, et al.
Pubblicazione: (2026)
DISTFLASHATTN: Distributed Memory-efficient Attention for Long-context LLMs Training
di: Li, Dacheng, et al.
Pubblicazione: (2023)
di: Li, Dacheng, et al.
Pubblicazione: (2023)
FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
di: Ye, Zihao, et al.
Pubblicazione: (2025)
di: Ye, Zihao, et al.
Pubblicazione: (2025)
Uncertainty-Aware Explainable Federated Learning
di: Zhang, Yanci, et al.
Pubblicazione: (2025)
di: Zhang, Yanci, et al.
Pubblicazione: (2025)
A-3PO: Accelerating Asynchronous LLM Training with Staleness-aware Proximal Policy Approximation
di: Li, Xiaocan, et al.
Pubblicazione: (2025)
di: Li, Xiaocan, et al.
Pubblicazione: (2025)
FT-Transformer: Resilient and Reliable Transformer with End-to-End Fault Tolerant Attention
di: Dai, Huangliang, et al.
Pubblicazione: (2025)
di: Dai, Huangliang, et al.
Pubblicazione: (2025)
Kascade: A Practical Sparse Attention Method for Long-Context LLM Inference
di: Deshmukh, Dhruv, et al.
Pubblicazione: (2025)
di: Deshmukh, Dhruv, et al.
Pubblicazione: (2025)
Topology-Aware Knowledge Propagation in Decentralized Learning
di: Sakarvadia, Mansi, et al.
Pubblicazione: (2025)
di: Sakarvadia, Mansi, et al.
Pubblicazione: (2025)
Federated Attention: A Distributed Paradigm for Collaborative LLM Inference over Edge Networks
di: Deng, Xiumei, et al.
Pubblicazione: (2025)
di: Deng, Xiumei, et al.
Pubblicazione: (2025)
MoSKA: Mixture of Shared KV Attention for Efficient Long-Sequence LLM Inference
di: Rhee, Myunghyun, et al.
Pubblicazione: (2025)
di: Rhee, Myunghyun, et al.
Pubblicazione: (2025)
CA-AFP: Cluster-Aware Adaptive Federated Pruning
di: Jha, Om Govind, et al.
Pubblicazione: (2026)
di: Jha, Om Govind, et al.
Pubblicazione: (2026)
RaMP: Runtime-Aware Megakernel Polymorphism for Mixture-of-Experts
di: Sharma, Vyom, et al.
Pubblicazione: (2026)
di: Sharma, Vyom, et al.
Pubblicazione: (2026)
Fairness-Aware Job Scheduling for Multi-Job Federated Learning
di: Shi, Yuxin, et al.
Pubblicazione: (2024)
di: Shi, Yuxin, et al.
Pubblicazione: (2024)
Game-Theoretic Deep Reinforcement Learning to Minimize Carbon Emissions and Energy Costs for AI Inference Workloads in Geo-Distributed Data Centers
di: Hogade, Ninad, et al.
Pubblicazione: (2024)
di: Hogade, Ninad, et al.
Pubblicazione: (2024)
TRAIL: Trust-Aware Client Scheduling for Semi-Decentralized Federated Learning
di: Hu, Gangqiang, et al.
Pubblicazione: (2024)
di: Hu, Gangqiang, et al.
Pubblicazione: (2024)
FlashMLA-ETAP: Efficient Transpose Attention Pipeline for Accelerating MLA Inference on NVIDIA H20 GPUs
di: Dege, Pengcuo, et al.
Pubblicazione: (2025)
di: Dege, Pengcuo, et al.
Pubblicazione: (2025)
Heterogeneity-Aware Resource Allocation and Topology Design for Hierarchical Federated Edge Learning
di: Gao, Zhidong, et al.
Pubblicazione: (2024)
di: Gao, Zhidong, et al.
Pubblicazione: (2024)
Low-Rank GEMM: Efficient Matrix Multiplication via Low-Rank Approximation with FP8 Acceleration
di: Metere, Alfredo
Pubblicazione: (2025)
di: Metere, Alfredo
Pubblicazione: (2025)
A deep cut into Split Federated Self-supervised Learning
di: Przewięźlikowski, Marcin, et al.
Pubblicazione: (2024)
di: Przewięźlikowski, Marcin, et al.
Pubblicazione: (2024)
SuperInfer: SLO-Aware Rotary Scheduling and Memory Management for LLM Inference on Superchips
di: Yu, Jiahuan, et al.
Pubblicazione: (2026)
di: Yu, Jiahuan, et al.
Pubblicazione: (2026)
CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters
di: Huang, Shaoyuan, et al.
Pubblicazione: (2026)
di: Huang, Shaoyuan, et al.
Pubblicazione: (2026)
Sustainable Carbon-Aware and Water-Efficient LLM Scheduling in Geo-Distributed Cloud Datacenters
di: Moore, Hayden, et al.
Pubblicazione: (2025)
di: Moore, Hayden, et al.
Pubblicazione: (2025)
LLM-PQ: Serving LLM on Heterogeneous Clusters with Phase-Aware Partition and Adaptive Quantization
di: Zhao, Juntao, et al.
Pubblicazione: (2024)
di: Zhao, Juntao, et al.
Pubblicazione: (2024)
A Framework for SLO, Carbon, and Wastewater-Aware Sustainable FaaS Cloud Platform Management
di: Qi, Sirui, et al.
Pubblicazione: (2024)
di: Qi, Sirui, et al.
Pubblicazione: (2024)
How Far Can Disaggregation Go? A Design-Space Exploration of Attention-FFN Disaggregation for Efficient MoE LLM Serving
di: Wu, Hanjiang, et al.
Pubblicazione: (2026)
di: Wu, Hanjiang, et al.
Pubblicazione: (2026)
Documenti analoghi
-
SATER: A Self-Aware and Token-Efficient Approach to Routing and Cascading
di: Shen, Yuanzhe, et al.
Pubblicazione: (2025) -
ATTENTION2D: Communication Efficient Distributed Self-Attention Mechanism
di: Elango, Venmugil
Pubblicazione: (2025) -
Cost-Efficient LLM Training with Lifetime-Aware Tensor Offloading via GPUDirect Storage
di: Yuan, Ziqi, et al.
Pubblicazione: (2025) -
Cost-Efficient Multimodal LLM Inference via Cross-Tier GPU Heterogeneity
di: Yu, Donglin
Pubblicazione: (2026) -
Token-Budget-Aware Pool Routing for Cost-Efficient LLM Inference
di: Chen, Huamin, et al.
Pubblicazione: (2026)