AsyncTLS: Efficient Generative LLM Inference with Asynchronous Two-level Sparse Attention
Fuente:
arXiv
Salvato in:
| Autori principali: | Hu, Yuxuan, Tan, Jianchao, Zhang, Jiaqi, Zan, Wen, Sun, Pingwei, Lu, Yifan, Sun, Yerui, Xie, Yuchen, Cai, Xunliang, Zhang, Jing |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Optimizing Native Sparse Attention with Latent Attention and Local Global Alternating Strategies
di: Hu, Yuxuan, et al.
Pubblicazione: (2025)
di: Hu, Yuxuan, et al.
Pubblicazione: (2025)
FG$^2$-GDN: Enhancing Long-Context Gated Delta Networks with Doubly Fine-Grained Control
di: Sun, Pingwei, et al.
Pubblicazione: (2026)
di: Sun, Pingwei, et al.
Pubblicazione: (2026)
SparseBalance: Load-Balanced Long Context Training with Dynamic Sparse Attention
di: Xu, Hongtao, et al.
Pubblicazione: (2026)
di: Xu, Hongtao, et al.
Pubblicazione: (2026)
Accelerate Speculative Decoding with Sparse Computation in Verification
di: Wang, Jikai, et al.
Pubblicazione: (2025)
di: Wang, Jikai, et al.
Pubblicazione: (2025)
MONA: Muon Optimizer with Nesterov Acceleration for Scalable Language Model Training
di: Li, Jiacheng, et al.
Pubblicazione: (2026)
di: Li, Jiacheng, et al.
Pubblicazione: (2026)
AFA-LoRA: Enabling Non-Linear Adaptations in LoRA with Activation Function Annealing
di: Li, Jiacheng, et al.
Pubblicazione: (2025)
di: Li, Jiacheng, et al.
Pubblicazione: (2025)
WISCA: A Lightweight Model Transition Method to Improve LLM Training via Weight Scaling
di: Li, Jiacheng, et al.
Pubblicazione: (2025)
di: Li, Jiacheng, et al.
Pubblicazione: (2025)
LoRS: Efficient Low-Rank Adaptation for Sparse Large Language Model
di: Hu, Yuxuan, et al.
Pubblicazione: (2025)
di: Hu, Yuxuan, et al.
Pubblicazione: (2025)
AsyncSparse: Accelerating Sparse Matrix-Matrix Multiplication on Asynchronous GPU Architectures
di: Liu, Jie, et al.
Pubblicazione: (2026)
di: Liu, Jie, et al.
Pubblicazione: (2026)
QUAD: Quantization and Parameter-Efficient Tuning of LLM with Activation Decomposition
di: Hu, Yuxuan, et al.
Pubblicazione: (2025)
di: Hu, Yuxuan, et al.
Pubblicazione: (2025)
C2T: A Classifier-Based Tree Construction Method in Speculative Decoding
di: Huo, Feiye, et al.
Pubblicazione: (2025)
di: Huo, Feiye, et al.
Pubblicazione: (2025)
How LLMs Are Persuaded: A Few Attention Heads, Rerouted
di: Sun, Xiangkun, et al.
Pubblicazione: (2026)
di: Sun, Xiangkun, et al.
Pubblicazione: (2026)
Mixture of Attention Spans: Optimizing LLM Inference Efficiency with Heterogeneous Sliding-Window Lengths
di: Fu, Tianyu, et al.
Pubblicazione: (2024)
di: Fu, Tianyu, et al.
Pubblicazione: (2024)
EyeLayer: Integrating Human Attention Patterns into LLM-Based Code Summarization
di: Zhang, Jiahao, et al.
Pubblicazione: (2026)
di: Zhang, Jiahao, et al.
Pubblicazione: (2026)
MaskPrune: Mask-based LLM Pruning for Layer-wise Uniform Structures
di: Qin, Jiayu, et al.
Pubblicazione: (2025)
di: Qin, Jiayu, et al.
Pubblicazione: (2025)
Resonant Context Anchoring: Decoupling Attention Routing and Signal Gain at Inference Time
di: Zhao, Mingkuan, et al.
Pubblicazione: (2026)
di: Zhao, Mingkuan, et al.
Pubblicazione: (2026)
AsyncFlow: An Asynchronous Streaming RL Framework for Efficient LLM Post-Training
di: Han, Zhenyu, et al.
Pubblicazione: (2025)
di: Han, Zhenyu, et al.
Pubblicazione: (2025)
AsyncSpade: Efficient Test-Time Scaling with Asynchronous Sparse Decoding
di: Luo, Shuqing, et al.
Pubblicazione: (2025)
di: Luo, Shuqing, et al.
Pubblicazione: (2025)
WeDLM: Reconciling Diffusion Language Models with Standard Causal Attention for Fast Inference
di: Liu, Aiwei, et al.
Pubblicazione: (2025)
di: Liu, Aiwei, et al.
Pubblicazione: (2025)
AsyncHZP: Hierarchical ZeRO Parallelism with Asynchronous Scheduling for Scalable LLM Training
di: Bai, Huawei, et al.
Pubblicazione: (2025)
di: Bai, Huawei, et al.
Pubblicazione: (2025)
AsyncDiff: Parallelizing Diffusion Models by Asynchronous Denoising
di: Chen, Zigeng, et al.
Pubblicazione: (2024)
di: Chen, Zigeng, et al.
Pubblicazione: (2024)
Structured-Sparse Attention for Entity Tracking with Subquadratic Sequence Complexity
di: Zhao, Hangyue, et al.
Pubblicazione: (2026)
di: Zhao, Hangyue, et al.
Pubblicazione: (2026)
SECURA: Sigmoid-Enhanced CUR Decomposition with Uninterrupted Retention and Low-Rank Adaptation in Large Language Models
di: Zhang, Yuxuan
Pubblicazione: (2025)
di: Zhang, Yuxuan
Pubblicazione: (2025)
POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference
di: Kamath, Aditya K, et al.
Pubblicazione: (2024)
di: Kamath, Aditya K, et al.
Pubblicazione: (2024)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
di: Ashuach, Tomer, et al.
Pubblicazione: (2025)
di: Ashuach, Tomer, et al.
Pubblicazione: (2025)
SAM Decoding: Speculative Decoding via Suffix Automaton
di: Hu, Yuxuan, et al.
Pubblicazione: (2024)
di: Hu, Yuxuan, et al.
Pubblicazione: (2024)
Can LLM Graph Reasoning Generalize beyond Pattern Memorization?
di: Zhang, Yizhuo, et al.
Pubblicazione: (2024)
di: Zhang, Yizhuo, et al.
Pubblicazione: (2024)
Async Control: Stress-testing Asynchronous Control Measures for LLM Agents
di: Stickland, Asa Cooper, et al.
Pubblicazione: (2025)
di: Stickland, Asa Cooper, et al.
Pubblicazione: (2025)
Solving the Granularity Mismatch: Hierarchical Preference Learning for Long-Horizon LLM Agents
di: Gao, Heyang, et al.
Pubblicazione: (2025)
di: Gao, Heyang, et al.
Pubblicazione: (2025)
Flash Communication: Reducing Tensor Parallelization Bottleneck for Fast Large Language Model Inference
di: Li, Qingyuan, et al.
Pubblicazione: (2024)
di: Li, Qingyuan, et al.
Pubblicazione: (2024)
TokenStack: A Heterogeneous HBM-PIM Architecture and Runtime for Efficient LLM Inference
di: Li, Zhuoran, et al.
Pubblicazione: (2026)
di: Li, Zhuoran, et al.
Pubblicazione: (2026)
AsyncDiff: Asynchronous Timestep Conditioning for Enhanced Text-to-Image Diffusion Inference
di: Xu, Longhuan, et al.
Pubblicazione: (2025)
di: Xu, Longhuan, et al.
Pubblicazione: (2025)
Understanding LLM Evaluator Behavior: A Structured Multi-Evaluator Framework for Merchant Risk Assessment
di: Wang, Liang, et al.
Pubblicazione: (2026)
di: Wang, Liang, et al.
Pubblicazione: (2026)
Streamlining Redundant Layers to Compress Large Language Models
di: Chen, Xiaodong, et al.
Pubblicazione: (2024)
di: Chen, Xiaodong, et al.
Pubblicazione: (2024)
Overcoming Long-Context Limitations of State-Space Models via Context-Dependent Sparse Attention
di: Zhan, Zhihao, et al.
Pubblicazione: (2025)
di: Zhan, Zhihao, et al.
Pubblicazione: (2025)
Fine-tuning vs Prompting, Can Language Models Understand Human Values?
di: Sun, Pingwei
Pubblicazione: (2024)
di: Sun, Pingwei
Pubblicazione: (2024)
SUBLLM: A Novel Efficient Architecture with Token Sequence Subsampling for LLM
di: Wang, Quandong, et al.
Pubblicazione: (2024)
di: Wang, Quandong, et al.
Pubblicazione: (2024)
Making Every Head Count: Sparse Attention Without the Speed-Performance Trade-off
di: Zhao, Mingkuan, et al.
Pubblicazione: (2025)
di: Zhao, Mingkuan, et al.
Pubblicazione: (2025)
Generalizable LLM Learning of Graph Synthetic Data with Post-training Alignment
di: Zhang, Yizhuo, et al.
Pubblicazione: (2025)
di: Zhang, Yizhuo, et al.
Pubblicazione: (2025)
EPS-MoE: Expert Pipeline Scheduler for Cost-Efficient MoE Inference
di: Qian, Yulei, et al.
Pubblicazione: (2024)
di: Qian, Yulei, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Optimizing Native Sparse Attention with Latent Attention and Local Global Alternating Strategies
di: Hu, Yuxuan, et al.
Pubblicazione: (2025) -
FG$^2$-GDN: Enhancing Long-Context Gated Delta Networks with Doubly Fine-Grained Control
di: Sun, Pingwei, et al.
Pubblicazione: (2026) -
SparseBalance: Load-Balanced Long Context Training with Dynamic Sparse Attention
di: Xu, Hongtao, et al.
Pubblicazione: (2026) -
Accelerate Speculative Decoding with Sparse Computation in Verification
di: Wang, Jikai, et al.
Pubblicazione: (2025) -
MONA: Muon Optimizer with Nesterov Acceleration for Scalable Language Model Training
di: Li, Jiacheng, et al.
Pubblicazione: (2026)