Bao, X., Xie, C., Tang, H., Weng, T., Wang, X., Zheng, Y., & Wang, X. (2025). DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding.
Chicago Style (17th ed.) CitationBao, Xiaoyi, Chenwei Xie, Hao Tang, Tingyu Weng, Xiaofeng Wang, Yun Zheng, and Xingang Wang. DynImg: Key Frames with Visual Prompts Are Good Representation for Multi-Modal Video Understanding. 2025.
MLA (9th ed.) CitationBao, Xiaoyi, et al. DynImg: Key Frames with Visual Prompts Are Good Representation for Multi-Modal Video Understanding. 2025.
Warning: These citations may not always be 100% accurate.