CLAA: Cross-Layer Attention Aggregation for Accelerating LLM Prefill
Fuente:
arXiv
Guardado en:
| Autores principales: | McDanel, Bradley, Li, Steven, Khaitan, Harshit |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
AMUSD: Asynchronous Multi-Device Speculative Decoding for LLM Acceleration
por: McDanel, Bradley
Publicado: (2024)
por: McDanel, Bradley
Publicado: (2024)
MoE-Spec: Expert Budgeting for Efficient Speculative Decoding
por: McDanel, Bradley, et al.
Publicado: (2026)
por: McDanel, Bradley, et al.
Publicado: (2026)
AttnCache: Accelerating Self-Attention Inference for LLM Prefill via Attention Cache
por: Song, Dinghong, et al.
Publicado: (2025)
por: Song, Dinghong, et al.
Publicado: (2025)
Block-Attention for Efficient Prefilling
por: Ma, Dongyang, et al.
Publicado: (2024)
por: Ma, Dongyang, et al.
Publicado: (2024)
LLM Router: Rethinking Routing with Prefill Activations
por: Varshney, Tanay, et al.
Publicado: (2026)
por: Varshney, Tanay, et al.
Publicado: (2026)
CritiPrefill: A Segment-wise Criticality-based Approach for Prefilling Acceleration in LLMs
por: Lv, Junlin, et al.
Publicado: (2024)
por: Lv, Junlin, et al.
Publicado: (2024)
IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse
por: Bai, Yushi, et al.
Publicado: (2026)
por: Bai, Yushi, et al.
Publicado: (2026)
Beyond Trusting Trust: Multi-Model Validation for Robust Code Generation
por: McDanel, Bradley
Publicado: (2025)
por: McDanel, Bradley
Publicado: (2025)
DSD: A Distributed Speculative Decoding Solution for Edge-Cloud Agile Large Model Serving
por: Yu, Fengze, et al.
Publicado: (2025)
por: Yu, Fengze, et al.
Publicado: (2025)
FastKV: Decoupling of Context Reduction and KV Cache Compression for Prefill-Decoding Acceleration
por: Jo, Dongwon, et al.
Publicado: (2025)
por: Jo, Dongwon, et al.
Publicado: (2025)
MoECollab: Democratizing LLM Development Through Collaborative Mixture of Experts
por: Harshit
Publicado: (2025)
por: Harshit
Publicado: (2025)
LIFT: Last-Mile Fine-Tuning for Table Explicitation
por: Khaitan, Divij, et al.
Publicado: (2026)
por: Khaitan, Divij, et al.
Publicado: (2026)
FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference
por: Lai, Xunhao, et al.
Publicado: (2025)
por: Lai, Xunhao, et al.
Publicado: (2025)
Accelerating Prefilling for Long-Context LLMs via Sparse Pattern Sharing
por: Peng, Dan, et al.
Publicado: (2025)
por: Peng, Dan, et al.
Publicado: (2025)
Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
por: Brandon, William, et al.
Publicado: (2024)
por: Brandon, William, et al.
Publicado: (2024)
Speculative Decoding and Beyond: An In-Depth Survey of Techniques
por: Hu, Yunhai, et al.
Publicado: (2025)
por: Hu, Yunhai, et al.
Publicado: (2025)
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
por: Liu, Di, et al.
Publicado: (2024)
por: Liu, Di, et al.
Publicado: (2024)
SkillAggregation: Reference-free LLM-Dependent Aggregation
por: Sun, Guangzhi, et al.
Publicado: (2024)
por: Sun, Guangzhi, et al.
Publicado: (2024)
Sockpuppetting: Jailbreaking LLMs by Combining Prefilling with Optimization
por: Dotsinski, Asen, et al.
Publicado: (2026)
por: Dotsinski, Asen, et al.
Publicado: (2026)
JacQuant: STE-Free Quantization-Aware Training via Learned Jacobian Surrogates
por: Yi, Kai, et al.
Publicado: (2026)
por: Yi, Kai, et al.
Publicado: (2026)
LayerBoost: Layer-Aware Attention Reduction for Efficient LLMs
por: Souibgui, Mohamed Ali, et al.
Publicado: (2026)
por: Souibgui, Mohamed Ali, et al.
Publicado: (2026)
Paying Attention to Facts: Quantifying the Knowledge Capacity of Attention Layers
por: Wong, Liang Ze
Publicado: (2025)
por: Wong, Liang Ze
Publicado: (2025)
High-Layer Attention Pruning with Rescaling
por: Liu, Songtao, et al.
Publicado: (2025)
por: Liu, Songtao, et al.
Publicado: (2025)
SqueezeAttention: 2D Management of KV-Cache in LLM Inference via Layer-wise Optimal Budget
por: Wang, Zihao, et al.
Publicado: (2024)
por: Wang, Zihao, et al.
Publicado: (2024)
SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention
por: Zhu, Qianchao, et al.
Publicado: (2024)
por: Zhu, Qianchao, et al.
Publicado: (2024)
R-Sparse: Rank-Aware Activation Sparsity for Efficient LLM Inference
por: Zhang, Zhenyu, et al.
Publicado: (2025)
por: Zhang, Zhenyu, et al.
Publicado: (2025)
Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse
por: Fu, Zizhuo, et al.
Publicado: (2026)
por: Fu, Zizhuo, et al.
Publicado: (2026)
Prefill-Guided Thinking for zero-shot detection of AI-generated images
por: Kachwala, Zoher, et al.
Publicado: (2025)
por: Kachwala, Zoher, et al.
Publicado: (2025)
Multi-Layer Attention is the Amplifier of Demonstration Effectiveness
por: Wang, Dingzirui, et al.
Publicado: (2025)
por: Wang, Dingzirui, et al.
Publicado: (2025)
Representation Consistency for Accurate and Coherent LLM Answer Aggregation
por: Jiang, Junqi, et al.
Publicado: (2025)
por: Jiang, Junqi, et al.
Publicado: (2025)
PipeSpec: Breaking Stage Dependencies in Hierarchical LLM Decoding
por: McDanel, Bradley, et al.
Publicado: (2025)
por: McDanel, Bradley, et al.
Publicado: (2025)
IntroLM: Introspective Language Models via Prefilling-Time Self-Evaluation
por: Kasnavieh, Hossein Hosseini, et al.
Publicado: (2026)
por: Kasnavieh, Hossein Hosseini, et al.
Publicado: (2026)
SwiftKV: Fast Prefill-Optimized Inference with Knowledge-Preserving Model Transformation
por: Qiao, Aurick, et al.
Publicado: (2024)
por: Qiao, Aurick, et al.
Publicado: (2024)
Human-LLM Hybrid Text Answer Aggregation for Crowd Annotations
por: Li, Jiyi
Publicado: (2024)
por: Li, Jiyi
Publicado: (2024)
Rethinking Attention: Exploring Shallow Feed-Forward Neural Networks as an Alternative to Attention Layers in Transformers
por: Bozic, Vukasin, et al.
Publicado: (2023)
por: Bozic, Vukasin, et al.
Publicado: (2023)
Enhancing LLM Evaluations: The Garbling Trick
por: Bradley, William F.
Publicado: (2024)
por: Bradley, William F.
Publicado: (2024)
MaxPoolBERT: Enhancing BERT Classification via Layer- and Token-Wise Aggregation
por: Behrendt, Maike, et al.
Publicado: (2025)
por: Behrendt, Maike, et al.
Publicado: (2025)
Mixture of Tokens: Continuous MoE through Cross-Example Aggregation
por: Antoniak, Szymon, et al.
Publicado: (2023)
por: Antoniak, Szymon, et al.
Publicado: (2023)
PreFT: Prefill-only finetuning for efficient inference
por: Lanpouthakoun, Andrew, et al.
Publicado: (2026)
por: Lanpouthakoun, Andrew, et al.
Publicado: (2026)
Mechanism and Emergence of Stacked Attention Heads in Multi-Layer Transformers
por: Musat, Tiberiu
Publicado: (2024)
por: Musat, Tiberiu
Publicado: (2024)
Ejemplares similares
-
AMUSD: Asynchronous Multi-Device Speculative Decoding for LLM Acceleration
por: McDanel, Bradley
Publicado: (2024) -
MoE-Spec: Expert Budgeting for Efficient Speculative Decoding
por: McDanel, Bradley, et al.
Publicado: (2026) -
AttnCache: Accelerating Self-Attention Inference for LLM Prefill via Attention Cache
por: Song, Dinghong, et al.
Publicado: (2025) -
Block-Attention for Efficient Prefilling
por: Ma, Dongyang, et al.
Publicado: (2024) -
LLM Router: Rethinking Routing with Prefill Activations
por: Varshney, Tanay, et al.
Publicado: (2026)