Guardado en:
| Autores principales: | Zhong, Wanli, Feng, Haibo, Zhou, Zirui, Peng, Hanyang, Yu, Shiqi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2511.21513 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models
por: Hu, Xing, et al.
Publicado: (2024)
por: Hu, Xing, et al.
Publicado: (2024)
Modality-Aware Zero-Shot Pruning and Sparse Attention for Efficient Multimodal Edge Inference
por: Sui, Yueyuan, et al.
Publicado: (2026)
por: Sui, Yueyuan, et al.
Publicado: (2026)
Unsupervised Multi-Attention Meta Transformer for Rotating Machinery Fault Diagnosis
por: Wang, Hanyang, et al.
Publicado: (2025)
por: Wang, Hanyang, et al.
Publicado: (2025)
AttentionEngine: A Versatile Framework for Efficient Attention Mechanisms on Diverse Hardware Platforms
por: Chen, Feiyang, et al.
Publicado: (2025)
por: Chen, Feiyang, et al.
Publicado: (2025)
Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference
por: Qiu, Quantong, et al.
Publicado: (2026)
por: Qiu, Quantong, et al.
Publicado: (2026)
SparQ Attention: Bandwidth-Efficient LLM Inference
por: Ribar, Luka, et al.
Publicado: (2023)
por: Ribar, Luka, et al.
Publicado: (2023)
Spatial Conformal Inference through Localized Quantile Regression
por: Jiang, Hanyang, et al.
Publicado: (2024)
por: Jiang, Hanyang, et al.
Publicado: (2024)
ScoutAttention: Efficient KV Cache Offloading via Layer-Ahead CPU Pre-computation for LLM Inference
por: Zhang, Qiuyang, et al.
Publicado: (2026)
por: Zhang, Qiuyang, et al.
Publicado: (2026)
ENA: Efficient N-dimensional Attention
por: Zhong, Yibo
Publicado: (2025)
por: Zhong, Yibo
Publicado: (2025)
CHAI: Clustered Head Attention for Efficient LLM Inference
por: Agarwal, Saurabh, et al.
Publicado: (2024)
por: Agarwal, Saurabh, et al.
Publicado: (2024)
Efficient Low Rank Attention for Long-Context Inference in Large Language Models
por: Li, Tenghui, et al.
Publicado: (2025)
por: Li, Tenghui, et al.
Publicado: (2025)
FlashMLA-ETAP: Efficient Transpose Attention Pipeline for Accelerating MLA Inference on NVIDIA H20 GPUs
por: Dege, Pengcuo, et al.
Publicado: (2025)
por: Dege, Pengcuo, et al.
Publicado: (2025)
Hermes: Memory-Efficient Pipeline Inference for Large Models on Edge Devices
por: Han, Xueyuan, et al.
Publicado: (2024)
por: Han, Xueyuan, et al.
Publicado: (2024)
Taming the Exponential: A Fast Softmax Surrogate for Integer-Native Edge Inference
por: Danopoulos, Dimitrios, et al.
Publicado: (2026)
por: Danopoulos, Dimitrios, et al.
Publicado: (2026)
Delta Attention: Fast and Accurate Sparse Attention Inference by Delta Correction
por: Willette, Jeffrey, et al.
Publicado: (2025)
por: Willette, Jeffrey, et al.
Publicado: (2025)
CSAttention: Centroid-Scoring Attention for Accelerating LLM Inference
por: Song, Chuxu, et al.
Publicado: (2026)
por: Song, Chuxu, et al.
Publicado: (2026)
AttnCache: Accelerating Self-Attention Inference for LLM Prefill via Attention Cache
por: Song, Dinghong, et al.
Publicado: (2025)
por: Song, Dinghong, et al.
Publicado: (2025)
Meta-Attention: Bayesian Per-Token Routing for Efficient Transformer Inference
por: Ferrari, Alan
Publicado: (2026)
por: Ferrari, Alan
Publicado: (2026)
FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference
por: Lai, Xunhao, et al.
Publicado: (2025)
por: Lai, Xunhao, et al.
Publicado: (2025)
INT-FlashAttention: Enabling Flash Attention for INT8 Quantization
por: Chen, Shimao, et al.
Publicado: (2024)
por: Chen, Shimao, et al.
Publicado: (2024)
Kernelized Edge Attention: Addressing Semantic Attention Blurring in Temporal Graph Neural Networks
por: Waghmare, Govind, et al.
Publicado: (2026)
por: Waghmare, Govind, et al.
Publicado: (2026)
NoMAD-Attention: Efficient LLM Inference on CPUs Through Multiply-add-free Attention
por: Zhang, Tianyi, et al.
Publicado: (2024)
por: Zhang, Tianyi, et al.
Publicado: (2024)
Spatial-Temporal Attention Model for Traffic State Estimation with Sparse Internet of Vehicles
por: Xue, Jianzhe, et al.
Publicado: (2024)
por: Xue, Jianzhe, et al.
Publicado: (2024)
SageAttention: Accurate 8-Bit Attention for Plug-and-play Inference Acceleration
por: Zhang, Jintao, et al.
Publicado: (2024)
por: Zhang, Jintao, et al.
Publicado: (2024)
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference
por: Deng, Weishu, et al.
Publicado: (2025)
por: Deng, Weishu, et al.
Publicado: (2025)
Attention Once Is All You Need: Efficient Streaming Inference with Stateful Transformers
por: Norgren, Victor
Publicado: (2026)
por: Norgren, Victor
Publicado: (2026)
DistrAttention: An Efficient and Flexible Self-Attention Mechanism on Modern GPUs
por: Jin, Haolin, et al.
Publicado: (2025)
por: Jin, Haolin, et al.
Publicado: (2025)
Federated Attention: A Distributed Paradigm for Collaborative LLM Inference over Edge Networks
por: Deng, Xiumei, et al.
Publicado: (2025)
por: Deng, Xiumei, et al.
Publicado: (2025)
Progressive Sparse Attention: Algorithm and System Co-design for Efficient Attention in LLM Serving
por: Zhou, Qihui, et al.
Publicado: (2025)
por: Zhou, Qihui, et al.
Publicado: (2025)
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference
por: Gong, Ping, et al.
Publicado: (2025)
por: Gong, Ping, et al.
Publicado: (2025)
Time-Aware Attention for Enhanced Electronic Health Records Modeling
por: Yu, Junhan, et al.
Publicado: (2025)
por: Yu, Junhan, et al.
Publicado: (2025)
ASAP: Attention-Shift-Aware Pruning for Efficient LVLM Inference
por: Pathak, Surendra, et al.
Publicado: (2026)
por: Pathak, Surendra, et al.
Publicado: (2026)
Star Attention: Efficient LLM Inference over Long Sequences
por: Acharya, Shantanu, et al.
Publicado: (2024)
por: Acharya, Shantanu, et al.
Publicado: (2024)
Symphony-MoE: Harmonizing Disparate Pre-trained Models into a Coherent Mixture-of-Experts
por: Wang, Qi, et al.
Publicado: (2025)
por: Wang, Qi, et al.
Publicado: (2025)
Edge Attention Module for Object Classification
por: Roy, Santanu, et al.
Publicado: (2025)
por: Roy, Santanu, et al.
Publicado: (2025)
Conv-Basis: A New Paradigm for Efficient Attention Inference and Gradient Computation in Transformers
por: Liang, Yingyu, et al.
Publicado: (2024)
por: Liang, Yingyu, et al.
Publicado: (2024)
Linear Attention Sequence Parallelism
por: Sun, Weigao, et al.
Publicado: (2024)
por: Sun, Weigao, et al.
Publicado: (2024)
HelixPipe: Efficient Distributed Training of Long Sequence Transformers with Attention Parallel Pipeline Parallelism
por: Zhang, Geng, et al.
Publicado: (2025)
por: Zhang, Geng, et al.
Publicado: (2025)
SSVEP-BiMA: Bifocal Masking Attention Leveraging Native and Symmetric-Antisymmetric Components for Robust SSVEP Decoding
por: Liu, Yuxin, et al.
Publicado: (2025)
por: Liu, Yuxin, et al.
Publicado: (2025)
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference
por: Tang, Xiaojuan, et al.
Publicado: (2025)
por: Tang, Xiaojuan, et al.
Publicado: (2025)
Ejemplares similares
-
I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models
por: Hu, Xing, et al.
Publicado: (2024) -
Modality-Aware Zero-Shot Pruning and Sparse Attention for Efficient Multimodal Edge Inference
por: Sui, Yueyuan, et al.
Publicado: (2026) -
Unsupervised Multi-Attention Meta Transformer for Rotating Machinery Fault Diagnosis
por: Wang, Hanyang, et al.
Publicado: (2025) -
AttentionEngine: A Versatile Framework for Efficient Attention Mechanisms on Diverse Hardware Platforms
por: Chen, Feiyang, et al.
Publicado: (2025) -
Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference
por: Qiu, Quantong, et al.
Publicado: (2026)