CHAI: Clustered Head Attention for Efficient LLM Inference
Fuente:
arXiv
Salvato in:
| Autori principali: | Agarwal, Saurabh, Acun, Bilge, Hosmer, Basil, Elhoushi, Mostafa, Lee, Yejin, Venkataraman, Shivaram, Papailiopoulos, Dimitris, Wu, Carole-Jean |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding
di: Elhoushi, Mostafa, et al.
Pubblicazione: (2024)
di: Elhoushi, Mostafa, et al.
Pubblicazione: (2024)
Decoding Speculative Decoding
di: Yan, Minghao, et al.
Pubblicazione: (2024)
di: Yan, Minghao, et al.
Pubblicazione: (2024)
Is Flash Attention Stable?
di: Golden, Alicia, et al.
Pubblicazione: (2024)
di: Golden, Alicia, et al.
Pubblicazione: (2024)
Scaling Inference-Efficient Language Models
di: Bian, Song, et al.
Pubblicazione: (2025)
di: Bian, Song, et al.
Pubblicazione: (2025)
SYMPHONY: Improving Memory Management for LLM Inference Workloads
di: Agarwal, Saurabh, et al.
Pubblicazione: (2024)
di: Agarwal, Saurabh, et al.
Pubblicazione: (2024)
Generative AI Beyond LLMs: System Implications of Multi-Modal Generation
di: Golden, Alicia, et al.
Pubblicazione: (2023)
di: Golden, Alicia, et al.
Pubblicazione: (2023)
Scalable LLM Reasoning Acceleration with Low-rank Distillation
di: Dong, Harry, et al.
Pubblicazione: (2025)
di: Dong, Harry, et al.
Pubblicazione: (2025)
Predictive Pipelined Decoding: A Compute-Latency Trade-off for Exact LLM Decoding
di: Yang, Seongjun, et al.
Pubblicazione: (2023)
di: Yang, Seongjun, et al.
Pubblicazione: (2023)
CATransformers: Carbon Aware Transformers Through Joint Model-Hardware Optimization
di: Wang, Irene, et al.
Pubblicazione: (2025)
di: Wang, Irene, et al.
Pubblicazione: (2025)
Calibrating Beyond English: Language Diversity for Better Quantized Multilingual LLM
di: Chimoto, Everlyn Asiko, et al.
Pubblicazione: (2026)
di: Chimoto, Everlyn Asiko, et al.
Pubblicazione: (2026)
Hybrid Architectures for Language Models: Systematic Analysis and Design Insights
di: Bae, Sangmin, et al.
Pubblicazione: (2025)
di: Bae, Sangmin, et al.
Pubblicazione: (2025)
Eva: Cost-Efficient Cloud-Based Cluster Scheduling
di: Chang, Tzu-Tao, et al.
Pubblicazione: (2025)
di: Chang, Tzu-Tao, et al.
Pubblicazione: (2025)
Beyond Efficiency: Scaling AI Sustainably
di: Wu, Carole-Jean, et al.
Pubblicazione: (2024)
di: Wu, Carole-Jean, et al.
Pubblicazione: (2024)
Guiding Giants: Lightweight Controllers for Weighted Activation Steering in LLMs
di: Hegazy, Amr, et al.
Pubblicazione: (2025)
di: Hegazy, Amr, et al.
Pubblicazione: (2025)
From Artificial Needles to Real Haystacks: Improving Retrieval Capabilities in LLMs by Finetuning on Synthetic Data
di: Xiong, Zheyang, et al.
Pubblicazione: (2024)
di: Xiong, Zheyang, et al.
Pubblicazione: (2024)
AST-T5: Structure-Aware Pretraining for Code Generation and Understanding
di: Gong, Linyuan, et al.
Pubblicazione: (2024)
di: Gong, Linyuan, et al.
Pubblicazione: (2024)
DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads
di: Xiao, Guangxuan, et al.
Pubblicazione: (2024)
di: Xiao, Guangxuan, et al.
Pubblicazione: (2024)
Gecko: An Efficient Neural Architecture Inherently Processing Sequences with Arbitrary Lengths
di: Ma, Xuezhe, et al.
Pubblicazione: (2026)
di: Ma, Xuezhe, et al.
Pubblicazione: (2026)
Tesserae: Scalable Placement Policies for Deep Learning Workloads
di: Bian, Song, et al.
Pubblicazione: (2025)
di: Bian, Song, et al.
Pubblicazione: (2025)
Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and Pitfalls
di: Kang, Feiyang, et al.
Pubblicazione: (2025)
di: Kang, Feiyang, et al.
Pubblicazione: (2025)
Characterizing and Efficiently Accelerating Multimodal Generation Model Inference
di: Lee, Yejin, et al.
Pubblicazione: (2024)
di: Lee, Yejin, et al.
Pubblicazione: (2024)
Extrapolation by Association: Length Generalization Transfer in Transformers
di: Cai, Ziyang, et al.
Pubblicazione: (2025)
di: Cai, Ziyang, et al.
Pubblicazione: (2025)
Unlocking the Potential of Renewable Energy Through Curtailment Prediction
di: Acun, Bilge, et al.
Pubblicazione: (2024)
di: Acun, Bilge, et al.
Pubblicazione: (2024)
Structure-Aware Fill-in-the-Middle Pretraining for Code
di: Gong, Linyuan, et al.
Pubblicazione: (2025)
di: Gong, Linyuan, et al.
Pubblicazione: (2025)
Endless Terminals: Scaling RL Environments for Terminal Agents
di: Gandhi, Kanishk, et al.
Pubblicazione: (2026)
di: Gandhi, Kanishk, et al.
Pubblicazione: (2026)
LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models
di: Chang, Tzu-Tao, et al.
Pubblicazione: (2025)
di: Chang, Tzu-Tao, et al.
Pubblicazione: (2025)
Evaluation of LLMs on Syntax-Aware Code Fill-in-the-Middle Tasks
di: Gong, Linyuan, et al.
Pubblicazione: (2024)
di: Gong, Linyuan, et al.
Pubblicazione: (2024)
Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs
di: Bian, Song, et al.
Pubblicazione: (2025)
di: Bian, Song, et al.
Pubblicazione: (2025)
ZigzagAttention: Efficient Long-Context Inference with Exclusive Retrieval and Streaming Heads
di: Liu, Zhuorui, et al.
Pubblicazione: (2025)
di: Liu, Zhuorui, et al.
Pubblicazione: (2025)
Brevity is the soul of wit: Pruning long files for code generation
di: Singh, Aaditya K., et al.
Pubblicazione: (2024)
di: Singh, Aaditya K., et al.
Pubblicazione: (2024)
CHAI: CacHe Attention Inference for text2video
di: Cherian, Joel Mathew, et al.
Pubblicazione: (2026)
di: Cherian, Joel Mathew, et al.
Pubblicazione: (2026)
PolyThrottle: Energy-efficient Neural Network Inference on Edge Devices
di: Yan, Minghao, et al.
Pubblicazione: (2023)
di: Yan, Minghao, et al.
Pubblicazione: (2023)
DyLLM: Efficient Diffusion LLM Inference via Saliency-based Token Selection and Partial Attention
di: Lee, Younjoo, et al.
Pubblicazione: (2026)
di: Lee, Younjoo, et al.
Pubblicazione: (2026)
Efficient LLM Inference with Kcache
di: He, Qiaozhi, et al.
Pubblicazione: (2024)
di: He, Qiaozhi, et al.
Pubblicazione: (2024)
Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
di: Shrivastava, Vaishnavi, et al.
Pubblicazione: (2025)
di: Shrivastava, Vaishnavi, et al.
Pubblicazione: (2025)
Efficient Long-Context LLM Inference via KV Cache Clustering
di: Hu, Jie, et al.
Pubblicazione: (2025)
di: Hu, Jie, et al.
Pubblicazione: (2025)
Attention Meets Reachability: Structural Equivalence and Efficiency in Grammar-Constrained LLM Decoding
di: Alpay, Faruk, et al.
Pubblicazione: (2026)
di: Alpay, Faruk, et al.
Pubblicazione: (2026)
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM
di: Fan, Zehao, et al.
Pubblicazione: (2025)
di: Fan, Zehao, et al.
Pubblicazione: (2025)
Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads
di: Cai, Tianle, et al.
Pubblicazione: (2024)
di: Cai, Tianle, et al.
Pubblicazione: (2024)
Everything Everywhere All at Once: LLMs can In-Context Learn Multiple Tasks in Superposition
di: Xiong, Zheyang, et al.
Pubblicazione: (2024)
di: Xiong, Zheyang, et al.
Pubblicazione: (2024)
Documenti analoghi
-
LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding
di: Elhoushi, Mostafa, et al.
Pubblicazione: (2024) -
Decoding Speculative Decoding
di: Yan, Minghao, et al.
Pubblicazione: (2024) -
Is Flash Attention Stable?
di: Golden, Alicia, et al.
Pubblicazione: (2024) -
Scaling Inference-Efficient Language Models
di: Bian, Song, et al.
Pubblicazione: (2025) -
SYMPHONY: Improving Memory Management for LLM Inference Workloads
di: Agarwal, Saurabh, et al.
Pubblicazione: (2024)