Bao, Z., Leng, J., Wang, J., Peng, B., & Lu, Y. (2026). Distilling Token-Trained Models into Byte-Level Models.
Chicago Style (17th ed.) CitationBao, Zishuo, Jiaqi Leng, Junxiong Wang, Bowen Peng, and Yucheng Lu. Distilling Token-Trained Models into Byte-Level Models. 2026.
MLA (9th ed.) CitationBao, Zishuo, et al. Distilling Token-Trained Models into Byte-Level Models. 2026.
Warning: These citations may not always be 100% accurate.