Block-Attention for Efficient Prefilling
Fuente:
arXiv
Salvato in:
| Autori principali: | Ma, Dongyang, Wang, Yan, Tian, Lan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
CritiPrefill: A Segment-wise Criticality-based Approach for Prefilling Acceleration in LLMs
di: Lv, Junlin, et al.
Pubblicazione: (2024)
di: Lv, Junlin, et al.
Pubblicazione: (2024)
Accelerating Prefilling for Long-Context LLMs via Sparse Pattern Sharing
di: Peng, Dan, et al.
Pubblicazione: (2025)
di: Peng, Dan, et al.
Pubblicazione: (2025)
Prefill-Guided Thinking for zero-shot detection of AI-generated images
di: Kachwala, Zoher, et al.
Pubblicazione: (2025)
di: Kachwala, Zoher, et al.
Pubblicazione: (2025)
MoBA: Mixture of Block Attention for Long-Context LLMs
di: Lu, Enzhe, et al.
Pubblicazione: (2025)
di: Lu, Enzhe, et al.
Pubblicazione: (2025)
SwiftKV: Fast Prefill-Optimized Inference with Knowledge-Preserving Model Transformation
di: Qiao, Aurick, et al.
Pubblicazione: (2024)
di: Qiao, Aurick, et al.
Pubblicazione: (2024)
IntroLM: Introspective Language Models via Prefilling-Time Self-Evaluation
di: Kasnavieh, Hossein Hosseini, et al.
Pubblicazione: (2026)
di: Kasnavieh, Hossein Hosseini, et al.
Pubblicazione: (2026)
Prepacking: A Simple Method for Fast Prefilling and Increased Throughput in Large Language Models
di: Zhao, Siyan, et al.
Pubblicazione: (2024)
di: Zhao, Siyan, et al.
Pubblicazione: (2024)
Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining
di: Fan, Dongyang, et al.
Pubblicazione: (2025)
di: Fan, Dongyang, et al.
Pubblicazione: (2025)
PreFT: Prefill-only finetuning for efficient inference
di: Lanpouthakoun, Andrew, et al.
Pubblicazione: (2026)
di: Lanpouthakoun, Andrew, et al.
Pubblicazione: (2026)
Towards Generalization of Block Attention via Automatic Segmentation and Block Distillation
di: Li, Shuaiyi, et al.
Pubblicazione: (2026)
di: Li, Shuaiyi, et al.
Pubblicazione: (2026)
FROST: Filtering Reasoning Outliers with Attention for Efficient Reasoning
di: Luo, Haozheng, et al.
Pubblicazione: (2026)
di: Luo, Haozheng, et al.
Pubblicazione: (2026)
Exposing the Systematic Vulnerability of Open-Weight Models to Prefill Attacks
di: Struppek, Lukas, et al.
Pubblicazione: (2026)
di: Struppek, Lukas, et al.
Pubblicazione: (2026)
Efficiently Dispatching Flash Attention For Partially Filled Attention Masks
di: Sharma, Agniv, et al.
Pubblicazione: (2024)
di: Sharma, Agniv, et al.
Pubblicazione: (2024)
Towards an empirical understanding of MoE design choices
di: Fan, Dongyang, et al.
Pubblicazione: (2024)
di: Fan, Dongyang, et al.
Pubblicazione: (2024)
Native Hybrid Attention for Efficient Sequence Modeling
di: Du, Jusen, et al.
Pubblicazione: (2025)
di: Du, Jusen, et al.
Pubblicazione: (2025)
Sliding Window Attention Training for Efficient Large Language Models
di: Fu, Zichuan, et al.
Pubblicazione: (2025)
di: Fu, Zichuan, et al.
Pubblicazione: (2025)
Attention Needs to Focus: A Unified Perspective on Attention Allocation
di: Fu, Zichuan, et al.
Pubblicazione: (2026)
di: Fu, Zichuan, et al.
Pubblicazione: (2026)
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid
di: Sun, Weigao, et al.
Pubblicazione: (2025)
di: Sun, Weigao, et al.
Pubblicazione: (2025)
Athena: Efficient Block-Wise Post-Training Quantization for Large Language Models Using Second-Order Matrix Derivative Information
di: Wang, Yanshu, et al.
Pubblicazione: (2024)
di: Wang, Yanshu, et al.
Pubblicazione: (2024)
NoMAD-Attention: Efficient LLM Inference on CPUs Through Multiply-add-free Attention
di: Zhang, Tianyi, et al.
Pubblicazione: (2024)
di: Zhang, Tianyi, et al.
Pubblicazione: (2024)
MiniCPM-SALA: Hybridizing Sparse and Linear Attention for Efficient Long-Context Modeling
di: MiniCPM Team, et al.
Pubblicazione: (2026)
di: MiniCPM Team, et al.
Pubblicazione: (2026)
Beyond KV Caching: Shared Attention for Efficient LLMs
di: Liao, Bingli, et al.
Pubblicazione: (2024)
di: Liao, Bingli, et al.
Pubblicazione: (2024)
JoMA: Demystifying Multilayer Transformers via JOint Dynamics of MLP and Attention
di: Tian, Yuandong, et al.
Pubblicazione: (2023)
di: Tian, Yuandong, et al.
Pubblicazione: (2023)
Scaling Stick-Breaking Attention: An Efficient Implementation and In-depth Study
di: Tan, Shawn, et al.
Pubblicazione: (2024)
di: Tan, Shawn, et al.
Pubblicazione: (2024)
Star Attention: Efficient LLM Inference over Long Sequences
di: Acharya, Shantanu, et al.
Pubblicazione: (2024)
di: Acharya, Shantanu, et al.
Pubblicazione: (2024)
More Expressive Attention with Negative Weights
di: Lv, Ang, et al.
Pubblicazione: (2024)
di: Lv, Ang, et al.
Pubblicazione: (2024)
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference
di: Tang, Xiaojuan, et al.
Pubblicazione: (2025)
di: Tang, Xiaojuan, et al.
Pubblicazione: (2025)
Bridging Internal Probability and Self-Consistency for Effective and Efficient LLM Reasoning
di: Zhou, Zhi, et al.
Pubblicazione: (2025)
di: Zhou, Zhi, et al.
Pubblicazione: (2025)
Efficient Prompt Tuning by Multi-Space Projection and Prompt Fusion
di: Lan, Pengxiang, et al.
Pubblicazione: (2024)
di: Lan, Pengxiang, et al.
Pubblicazione: (2024)
Taipan: Efficient and Expressive State Space Language Models with Selective Attention
di: Van Nguyen, Chien, et al.
Pubblicazione: (2024)
di: Van Nguyen, Chien, et al.
Pubblicazione: (2024)
Training Tensor Attention Efficiently: From Cubic to Almost Linear Time
di: Cao, Yang, et al.
Pubblicazione: (2024)
di: Cao, Yang, et al.
Pubblicazione: (2024)
Every Attention Matters: An Efficient Hybrid Architecture for Long-Context Reasoning
di: Ling Team, et al.
Pubblicazione: (2025)
di: Ling Team, et al.
Pubblicazione: (2025)
One-Pass to Reason: Token Duplication and Block-Sparse Mask for Efficient Fine-Tuning on Multi-Turn Reasoning
di: Goru, Ritesh, et al.
Pubblicazione: (2025)
di: Goru, Ritesh, et al.
Pubblicazione: (2025)
Hybrid Linear Attention Done Right: Efficient Distillation and Effective Architectures for Extremely Long Contexts
di: Chen, Yingfa, et al.
Pubblicazione: (2026)
di: Chen, Yingfa, et al.
Pubblicazione: (2026)
When Linear Attention Meets Autoregressive Decoding: Towards More Effective and Efficient Linearized Large Language Models
di: You, Haoran, et al.
Pubblicazione: (2024)
di: You, Haoran, et al.
Pubblicazione: (2024)
Scalable Efficient Training of Large Language Models with Low-dimensional Projected Attention
di: Lv, Xingtai, et al.
Pubblicazione: (2024)
di: Lv, Xingtai, et al.
Pubblicazione: (2024)
FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference
di: Lai, Xunhao, et al.
Pubblicazione: (2025)
di: Lai, Xunhao, et al.
Pubblicazione: (2025)
LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference
di: Wang, Guangtao, et al.
Pubblicazione: (2025)
di: Wang, Guangtao, et al.
Pubblicazione: (2025)
AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs
di: S, Santhosh G, et al.
Pubblicazione: (2025)
di: S, Santhosh G, et al.
Pubblicazione: (2025)
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
di: Yuan, Jingyang, et al.
Pubblicazione: (2025)
di: Yuan, Jingyang, et al.
Pubblicazione: (2025)
Documenti analoghi
-
CritiPrefill: A Segment-wise Criticality-based Approach for Prefilling Acceleration in LLMs
di: Lv, Junlin, et al.
Pubblicazione: (2024) -
Accelerating Prefilling for Long-Context LLMs via Sparse Pattern Sharing
di: Peng, Dan, et al.
Pubblicazione: (2025) -
Prefill-Guided Thinking for zero-shot detection of AI-generated images
di: Kachwala, Zoher, et al.
Pubblicazione: (2025) -
MoBA: Mixture of Block Attention for Long-Context LLMs
di: Lu, Enzhe, et al.
Pubblicazione: (2025) -
SwiftKV: Fast Prefill-Optimized Inference with Knowledge-Preserving Model Transformation
di: Qiao, Aurick, et al.
Pubblicazione: (2024)