Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters
Fuente:
arXiv
Guardado en:
| Autores principales: | Shyam, Vasudev, Pilault, Jonathan, Shepperd, Emily, Anthony, Quentin, Millidge, Beren |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
The Zamba2 Suite: Technical Report
por: Glorioso, Paolo, et al.
Publicado: (2024)
por: Glorioso, Paolo, et al.
Publicado: (2024)
Zamba: A Compact 7B SSM Hybrid Model
por: Glorioso, Paolo, et al.
Publicado: (2024)
por: Glorioso, Paolo, et al.
Publicado: (2024)
Toward Conversational Agents with Context and Time Sensitive Long-term Memory
por: Alonso, Nick, et al.
Publicado: (2024)
por: Alonso, Nick, et al.
Publicado: (2024)
BlackMamba: Mixture of Experts for State-Space Models
por: Anthony, Quentin, et al.
Publicado: (2024)
por: Anthony, Quentin, et al.
Publicado: (2024)
Compressed Convolutional Attention: Efficient Attention in a Compressed Latent Space
por: Figliolia, Tomas, et al.
Publicado: (2025)
por: Figliolia, Tomas, et al.
Publicado: (2025)
Zyda: A 1.3T Dataset for Open Language Modeling
por: Tokpanov, Yury, et al.
Publicado: (2024)
por: Tokpanov, Yury, et al.
Publicado: (2024)
Generalising E-prop to Deep Networks
por: Millidge, Beren
Publicado: (2025)
por: Millidge, Beren
Publicado: (2025)
Online Vector Quantized Attention
por: Alonso, Nick, et al.
Publicado: (2026)
por: Alonso, Nick, et al.
Publicado: (2026)
Mixture-of-PageRanks: Replacing Long-Context with Real-Time, Sparse GraphRAG
por: Alonso, Nicholas, et al.
Publicado: (2024)
por: Alonso, Nicholas, et al.
Publicado: (2024)
Zyda-2: a 5 Trillion Token High-Quality Dataset
por: Tokpanov, Yury, et al.
Publicado: (2024)
por: Tokpanov, Yury, et al.
Publicado: (2024)
Multipole Attention for Efficient Long Context Reasoning
por: Hooper, Coleman, et al.
Publicado: (2025)
por: Hooper, Coleman, et al.
Publicado: (2025)
S$^3$-Attention:Attention-Aligned Endogenous Retrieval for Memory-Bounded Long-Context Inference
por: Ma, Qingsen, et al.
Publicado: (2026)
por: Ma, Qingsen, et al.
Publicado: (2026)
Dynamic Attention-Guided Context Decoding for Mitigating Context Faithfulness Hallucinations in Large Language Models
por: Huang, Yanwen, et al.
Publicado: (2025)
por: Huang, Yanwen, et al.
Publicado: (2025)
LLMSteer: Improving Long-Context LLM Inference by Steering Attention on Reused Contexts
por: Gu, Zhuohan, et al.
Publicado: (2024)
por: Gu, Zhuohan, et al.
Publicado: (2024)
Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification
por: Donhauser, Konstantin, et al.
Publicado: (2025)
por: Donhauser, Konstantin, et al.
Publicado: (2025)
Hardware-Efficient Attention for Fast Decoding
por: Zadouri, Ted, et al.
Publicado: (2025)
por: Zadouri, Ted, et al.
Publicado: (2025)
SEAL: Scaling to Emphasize Attention for Long-Context Retrieval
por: Lee, Changhun, et al.
Publicado: (2025)
por: Lee, Changhun, et al.
Publicado: (2025)
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
por: Liu, Di, et al.
Publicado: (2024)
por: Liu, Di, et al.
Publicado: (2024)
DELTA: Dynamic Layer-Aware Token Attention for Efficient Long-Context Reasoning
por: Zarch, Hossein Entezari, et al.
Publicado: (2025)
por: Zarch, Hossein Entezari, et al.
Publicado: (2025)
Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection
por: Jo, Dongwon, et al.
Publicado: (2026)
por: Jo, Dongwon, et al.
Publicado: (2026)
Transformer Based Linear Attention with Optimized GPU Kernel Implementation
por: Gerami, Armin, et al.
Publicado: (2025)
por: Gerami, Armin, et al.
Publicado: (2025)
Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference
por: Qiu, Quantong, et al.
Publicado: (2026)
por: Qiu, Quantong, et al.
Publicado: (2026)
Analyzing the Attention Heads for Pronoun Disambiguation in Context-aware Machine Translation Models
por: Mąka, Paweł, et al.
Publicado: (2024)
por: Mąka, Paweł, et al.
Publicado: (2024)
Breaking Quadratic Barriers: A Non-Attention LLM for Ultra-Long Context Horizons
por: Kiruluta, Andrew, et al.
Publicado: (2025)
por: Kiruluta, Andrew, et al.
Publicado: (2025)
SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention
por: Zhu, Qianchao, et al.
Publicado: (2024)
por: Zhu, Qianchao, et al.
Publicado: (2024)
HiCI: Hierarchical Construction-Integration for Long-Context Attention
por: Zeng, Xiangyu, et al.
Publicado: (2026)
por: Zeng, Xiangyu, et al.
Publicado: (2026)
Cost-Optimal Grouped-Query Attention for Long-Context Modeling
por: Chen, Yingfa, et al.
Publicado: (2025)
por: Chen, Yingfa, et al.
Publicado: (2025)
MoBA: Mixture of Block Attention for Long-Context LLMs
por: Lu, Enzhe, et al.
Publicado: (2025)
por: Lu, Enzhe, et al.
Publicado: (2025)
Mixture of Attentions For Speculative Decoding
por: Zimmer, Matthieu, et al.
Publicado: (2024)
por: Zimmer, Matthieu, et al.
Publicado: (2024)
EntmaxKV: Support-Aware Decoding for Entmax Attention
por: Duarte, Gonçalo, et al.
Publicado: (2026)
por: Duarte, Gonçalo, et al.
Publicado: (2026)
FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference
por: Lai, Xunhao, et al.
Publicado: (2025)
por: Lai, Xunhao, et al.
Publicado: (2025)
Every Attention Matters: An Efficient Hybrid Architecture for Long-Context Reasoning
por: Ling Team, et al.
Publicado: (2025)
por: Ling Team, et al.
Publicado: (2025)
Attention-aware semantic relevance predicting Chinese sentence reading
por: Sun, Kun
Publicado: (2024)
por: Sun, Kun
Publicado: (2024)
Token Distillation: Attention-aware Input Embeddings For New Tokens
por: Dobler, Konstantin, et al.
Publicado: (2025)
por: Dobler, Konstantin, et al.
Publicado: (2025)
Faster Transformer Decoding: N-gram Masked Self-Attention
por: Chelba, Ciprian, et al.
Publicado: (2020)
por: Chelba, Ciprian, et al.
Publicado: (2020)
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM
por: Fan, Zehao, et al.
Publicado: (2025)
por: Fan, Zehao, et al.
Publicado: (2025)
MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
por: Jiang, Huiqiang, et al.
Publicado: (2024)
por: Jiang, Huiqiang, et al.
Publicado: (2024)
Ltri-LLM: Streaming Long Context Inference for LLMs with Training-Free Dynamic Triangular Attention Pattern
por: Tang, Hongyin, et al.
Publicado: (2024)
por: Tang, Hongyin, et al.
Publicado: (2024)
LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference
por: Wang, Guangtao, et al.
Publicado: (2025)
por: Wang, Guangtao, et al.
Publicado: (2025)
Found in the Middle: Calibrating Positional Attention Bias Improves Long Context Utilization
por: Hsieh, Cheng-Yu, et al.
Publicado: (2024)
por: Hsieh, Cheng-Yu, et al.
Publicado: (2024)
Ejemplares similares
-
The Zamba2 Suite: Technical Report
por: Glorioso, Paolo, et al.
Publicado: (2024) -
Zamba: A Compact 7B SSM Hybrid Model
por: Glorioso, Paolo, et al.
Publicado: (2024) -
Toward Conversational Agents with Context and Time Sensitive Long-term Memory
por: Alonso, Nick, et al.
Publicado: (2024) -
BlackMamba: Mixture of Experts for State-Space Models
por: Anthony, Quentin, et al.
Publicado: (2024) -
Compressed Convolutional Attention: Efficient Attention in a Compressed Latent Space
por: Figliolia, Tomas, et al.
Publicado: (2025)