Hsieh, T., Choi, H., & Kim, M. (2024). Multimodal Representation Loss Between Timed Text and Audio for Regularized Speech Separation.
Chicago Style (17th ed.) CitationHsieh, Tsun-An, Heeyoul Choi, and Minje Kim. Multimodal Representation Loss Between Timed Text and Audio for Regularized Speech Separation. 2024.
MLA (9th ed.) CitationHsieh, Tsun-An, et al. Multimodal Representation Loss Between Timed Text and Audio for Regularized Speech Separation. 2024.
Warning: These citations may not always be 100% accurate.