Ning, Z., Zhao, J., Jin, Q., Ding, W., & Guo, M. (2024). Inf-MLLM: Efficient Streaming Inference of Multimodal Large Language Models on a Single GPU.
Cita Chicago Style (17a ed.)Ning, Zhenyu, Jieru Zhao, Qihao Jin, Wenchao Ding, y Minyi Guo. Inf-MLLM: Efficient Streaming Inference of Multimodal Large Language Models on a Single GPU. 2024.
Cita MLA (9a ed.)Ning, Zhenyu, et al. Inf-MLLM: Efficient Streaming Inference of Multimodal Large Language Models on a Single GPU. 2024.
Precaución: Estas citas no son 100% exactas.