Wang, Z., Cui, B., & Gan, S. (2024). SqueezeAttention: 2D Management of KV-Cache in LLM Inference via Layer-wise Optimal Budget.
Citazione stile Chigago Style (17a edizione)Wang, Zihao, Bin Cui, e Shaoduo Gan. SqueezeAttention: 2D Management of KV-Cache in LLM Inference via Layer-wise Optimal Budget. 2024.
Citatione MLA (9a ed.)Wang, Zihao, et al. SqueezeAttention: 2D Management of KV-Cache in LLM Inference via Layer-wise Optimal Budget. 2024.
Attenzione: Queste citazioni potrebbero non essere precise al 100%.