Gim, I., Chen, G., Lee, S., Sarda, N., Khandelwal, A., & Zhong, L. (2023). Prompt Cache: Modular Attention Reuse for Low-Latency Inference.
Chicago Style (17th ed.) CitationGim, In, Guojun Chen, Seung-seob Lee, Nikhil Sarda, Anurag Khandelwal, and Lin Zhong. Prompt Cache: Modular Attention Reuse for Low-Latency Inference. 2023.
MLA (9th ed.) CitationGim, In, et al. Prompt Cache: Modular Attention Reuse for Low-Latency Inference. 2023.
Warning: These citations may not always be 100% accurate.