Li, Z., Zhang, X., Guo, Y., Bennamoun, M., Boussaid, F., Dwivedi, G., . . . Ke, Q. (2025). Watch and Listen: Understanding Audio-Visual-Speech Moments with Multimodal LLM.
Chicago Style (17th ed.) CitationLi, Zinuo, Xian Zhang, Yongxin Guo, Mohammed Bennamoun, Farid Boussaid, Girish Dwivedi, Luqi Gong, and Qiuhong Ke. Watch and Listen: Understanding Audio-Visual-Speech Moments with Multimodal LLM. 2025.
MLA (9th ed.) CitationLi, Zinuo, et al. Watch and Listen: Understanding Audio-Visual-Speech Moments with Multimodal LLM. 2025.
Warning: These citations may not always be 100% accurate.