PAT: Accelerating LLM Decoding via Prefix-Aware Attention with Resource Efficient Multi-Tile Kernel

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yi, Jinjun, Zhao, Zhixin, Hu, Yitao, Yan, Ke, Sun, Weiwei, Wang, Hao, Zhao, Laiping, Zhang, Yuhao, Li, Wenxin, Li, Keqiu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!

Similar Items