Jiang, D., Liu, Y., Liu, S., Zhao, J., Zhang, H., Gao, Z., . . . Xiong, H. (2023). From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models.
Chicago-Zitierstil (17. Ausg.)Jiang, Dongsheng, Yuchen Liu, Songlin Liu, Jin'e Zhao, Hao Zhang, Zhen Gao, Xiaopeng Zhang, Jin Li, und Hongkai Xiong. From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models. 2023.
MLA-Zitierstil (9. Ausg.)Jiang, Dongsheng, et al. From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models. 2023.
Achtung: Diese Zitate sind unter Umständen nicht zu 100% korrekt.