CSAttention: Centroid-Scoring Attention for Accelerating LLM Inference
Fuente:
arXiv
Guardado en:
| Autores principales: | Song, Chuxu, Peng, Zhencan, Wei, Jiuqi, Yang, Chuanhui |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Almost Sure Convergence of Linear Temporal Difference Learning with Arbitrary Features
por: Wang, Jiuqi, et al.
Publicado: (2024)
por: Wang, Jiuqi, et al.
Publicado: (2024)
SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention
por: Zhu, Qianchao, et al.
Publicado: (2024)
por: Zhu, Qianchao, et al.
Publicado: (2024)
Almost Sure Convergence of Differential Temporal Difference Learning for Average Reward Markov Decision Processes
por: Blaser, Ethan, et al.
Publicado: (2026)
por: Blaser, Ethan, et al.
Publicado: (2026)
Same Signal, Opposite Meaning: Direction-Informed Adaptive Learning for LLM Agents
por: Li, Ziming, et al.
Publicado: (2026)
por: Li, Ziming, et al.
Publicado: (2026)
Generative Risk Minimization for Out-of-Distribution Generalization on Graphs
por: Wang, Song, et al.
Publicado: (2025)
por: Wang, Song, et al.
Publicado: (2025)
LLM-Empowered Class Imbalanced Graph Prompt Learning for Online Drug Trafficking Detection
por: Ma, Tianyi, et al.
Publicado: (2025)
por: Ma, Tianyi, et al.
Publicado: (2025)
Scout Before You Attend: Sketch-and-Walk Sparse Attention for Efficient LLM Inference
por: Le, Hoang Anh Duy, et al.
Publicado: (2026)
por: Le, Hoang Anh Duy, et al.
Publicado: (2026)
Breaking the Reclustering Barrier in Centroid-based Deep Clustering
por: Miklautz, Lukas, et al.
Publicado: (2024)
por: Miklautz, Lukas, et al.
Publicado: (2024)
Accelerating LLM Inference Throughput via Asynchronous KV Cache Prefetching
por: Dong, Yanhao, et al.
Publicado: (2025)
por: Dong, Yanhao, et al.
Publicado: (2025)
FastMTP: Accelerating LLM Inference with Enhanced Multi-Token Prediction
por: Cai, Yuxuan, et al.
Publicado: (2025)
por: Cai, Yuxuan, et al.
Publicado: (2025)
MISA: Mixture of Indexer Sparse Attention for Long-Context LLM Inference
por: Zhou, Ruijie, et al.
Publicado: (2026)
por: Zhou, Ruijie, et al.
Publicado: (2026)
Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration
por: Wen, Zhuofan, et al.
Publicado: (2024)
por: Wen, Zhuofan, et al.
Publicado: (2024)
Learning with Noisy Labels through Learnable Weighting and Centroid Similarity
por: Wani, Farooq Ahmad, et al.
Publicado: (2023)
por: Wani, Farooq Ahmad, et al.
Publicado: (2023)
CATS: Cascaded Adaptive Tree Speculation for Memory-Limited LLM Inference Acceleration
por: Han, Yuning, et al.
Publicado: (2026)
por: Han, Yuning, et al.
Publicado: (2026)
Recursive Speculative Decoding: Accelerating LLM Inference via Sampling Without Replacement
por: Jeon, Wonseok, et al.
Publicado: (2024)
por: Jeon, Wonseok, et al.
Publicado: (2024)
Multi-LLM Adaptive Conformal Inference for Reliable LLM Responses
por: Noh, Kangjun, et al.
Publicado: (2026)
por: Noh, Kangjun, et al.
Publicado: (2026)
Quaternion Self-Attention with Shared Scores
por: Yamauchi, Shogo, et al.
Publicado: (2026)
por: Yamauchi, Shogo, et al.
Publicado: (2026)
Generative Score Inference for Multimodal Data
por: Tian, Xinyu, et al.
Publicado: (2026)
por: Tian, Xinyu, et al.
Publicado: (2026)
Evaluating LLM Safety Under Repeated Inference via Accelerated Prompt Stress Testing
por: Broadwater, Keita
Publicado: (2026)
por: Broadwater, Keita
Publicado: (2026)
HashAttention: Semantic Sparsity for Faster Inference
por: Desai, Aditya, et al.
Publicado: (2024)
por: Desai, Aditya, et al.
Publicado: (2024)
Hyperparameters in Score-Based Membership Inference Attacks
por: Pradhan, Gauri, et al.
Publicado: (2025)
por: Pradhan, Gauri, et al.
Publicado: (2025)
Gradient-Direction Sensitivity Reveals Linear-Centroid Coupling Hidden by Optimizer Trajectories
por: Xu, Yongzhong
Publicado: (2026)
por: Xu, Yongzhong
Publicado: (2026)
Generalizing Behavior via Inverse Reinforcement Learning with Closed-Form Reward Centroids
por: Lazzati, Filippo, et al.
Publicado: (2025)
por: Lazzati, Filippo, et al.
Publicado: (2025)
Is One Score Enough? Rethinking the Evaluation of Sequentially Evolving LLM Memory
por: Dong, Songwei, et al.
Publicado: (2026)
por: Dong, Songwei, et al.
Publicado: (2026)
Experience Replay Addresses Loss of Plasticity in Continual Learning
por: Wang, Jiuqi, et al.
Publicado: (2025)
por: Wang, Jiuqi, et al.
Publicado: (2025)
Bregman Centroid Guided Cross-Entropy Method
por: Gu, Yuliang, et al.
Publicado: (2025)
por: Gu, Yuliang, et al.
Publicado: (2025)
HSR-Enhanced Sparse Attention Acceleration
por: Chen, Bo, et al.
Publicado: (2024)
por: Chen, Bo, et al.
Publicado: (2024)
SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model Inference
por: Zhang, Jintao, et al.
Publicado: (2025)
por: Zhang, Jintao, et al.
Publicado: (2025)
Star Attention: Efficient LLM Inference over Long Sequences
por: Acharya, Shantanu, et al.
Publicado: (2024)
por: Acharya, Shantanu, et al.
Publicado: (2024)
Controllable Graph Generation with Diffusion Models via Inference-Time Tree Search Guidance
por: Zhao, Jiachi, et al.
Publicado: (2025)
por: Zhao, Jiachi, et al.
Publicado: (2025)
RAP: Runtime Adaptive Pruning for LLM Inference
por: Liu, Huanrong, et al.
Publicado: (2025)
por: Liu, Huanrong, et al.
Publicado: (2025)
Entropy Centroids as Intrinsic Rewards for Test-Time Scaling
por: Zhao, Wenshuo, et al.
Publicado: (2026)
por: Zhao, Wenshuo, et al.
Publicado: (2026)
NoMAD-Attention: Efficient LLM Inference on CPUs Through Multiply-add-free Attention
por: Zhang, Tianyi, et al.
Publicado: (2024)
por: Zhang, Tianyi, et al.
Publicado: (2024)
BuddyMoE: Exploiting Expert Redundancy to Accelerate Memory-Constrained Mixture-of-Experts Inference
por: Wang, Yun, et al.
Publicado: (2025)
por: Wang, Yun, et al.
Publicado: (2025)
TIDE: Temporal Incremental Draft Engine for Self-Improving LLM Inference
por: Park, Jiyoung, et al.
Publicado: (2026)
por: Park, Jiyoung, et al.
Publicado: (2026)
Accelerating LLM Inference with Lossless Speculative Decoding Algorithms for Heterogeneous Vocabularies
por: Timor, Nadav, et al.
Publicado: (2025)
por: Timor, Nadav, et al.
Publicado: (2025)
A Multi-Dimensional Quality Scoring Framework for Decentralized LLM Inference with Proof of Quality
por: Tian, Arther, et al.
Publicado: (2026)
por: Tian, Arther, et al.
Publicado: (2026)
Ragged Paged Attention: A High-Performance and Flexible LLM Inference Kernel for TPU
por: Jiang, Jevin, et al.
Publicado: (2026)
por: Jiang, Jevin, et al.
Publicado: (2026)
Conv-Basis: A New Paradigm for Efficient Attention Inference and Gradient Computation in Transformers
por: Liang, Yingyu, et al.
Publicado: (2024)
por: Liang, Yingyu, et al.
Publicado: (2024)
AccLLM: Accelerating Long-Context LLM Inference Via Algorithm-Hardware Co-Design
por: Liang, Yanbiao, et al.
Publicado: (2025)
por: Liang, Yanbiao, et al.
Publicado: (2025)
Ejemplares similares
-
Almost Sure Convergence of Linear Temporal Difference Learning with Arbitrary Features
por: Wang, Jiuqi, et al.
Publicado: (2024) -
SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention
por: Zhu, Qianchao, et al.
Publicado: (2024) -
Almost Sure Convergence of Differential Temporal Difference Learning for Average Reward Markov Decision Processes
por: Blaser, Ethan, et al.
Publicado: (2026) -
Same Signal, Opposite Meaning: Direction-Informed Adaptive Learning for LLM Agents
por: Li, Ziming, et al.
Publicado: (2026) -
Generative Risk Minimization for Out-of-Distribution Generalization on Graphs
por: Wang, Song, et al.
Publicado: (2025)