Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Liang, Wang, Zekun, Ren, Shuhuai, Li, Lei, Zhao, Haozhe, Li, Yunshui, Cai, Zefan, Guo, Hongcheng, Zhang, Lei, Xiong, Yizhe, Zhang, Yichi, Wu, Ruoyu, Dong, Qingxiu, Zhang, Ge, Yang, Jian, Meng, Lingwei, Hu, Shujie, Chen, Yulong, Lin, Junyang, Bai, Shuai, Vlachos, Andreas, Tan, Xu, Zhang, Minjia, Xiao, Wen, Yee, Aaron, Liu, Tianyu, Chang, Baobao
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915083438784512
author Chen, Liang
Wang, Zekun
Ren, Shuhuai
Li, Lei
Zhao, Haozhe
Li, Yunshui
Cai, Zefan
Guo, Hongcheng
Zhang, Lei
Xiong, Yizhe
Zhang, Yichi
Wu, Ruoyu
Dong, Qingxiu
Zhang, Ge
Yang, Jian
Meng, Lingwei
Hu, Shujie
Chen, Yulong
Lin, Junyang
Bai, Shuai
Vlachos, Andreas
Tan, Xu
Zhang, Minjia
Xiao, Wen
Yee, Aaron
Liu, Tianyu
Chang, Baobao
author_facet Chen, Liang
Wang, Zekun
Ren, Shuhuai
Li, Lei
Zhao, Haozhe
Li, Yunshui
Cai, Zefan
Guo, Hongcheng
Zhang, Lei
Xiong, Yizhe
Zhang, Yichi
Wu, Ruoyu
Dong, Qingxiu
Zhang, Ge
Yang, Jian
Meng, Lingwei
Hu, Shujie
Chen, Yulong
Lin, Junyang
Bai, Shuai
Vlachos, Andreas
Tan, Xu
Zhang, Minjia
Xiao, Wen
Yee, Aaron
Liu, Tianyu
Chang, Baobao
contents Building on the foundations of language modeling in natural language processing, Next Token Prediction (NTP) has evolved into a versatile training objective for machine learning tasks across various modalities, achieving considerable success. As Large Language Models (LLMs) have advanced to unify understanding and generation tasks within the textual modality, recent research has shown that tasks from different modalities can also be effectively encapsulated within the NTP framework, transforming the multimodal information into tokens and predict the next one given the context. This survey introduces a comprehensive taxonomy that unifies both understanding and generation within multimodal learning through the lens of NTP. The proposed taxonomy covers five key aspects: Multimodal tokenization, MMNTP model architectures, unified task representation, datasets \& evaluation, and open challenges. This new taxonomy aims to aid researchers in their exploration of multimodal intelligence. An associated GitHub repository collecting the latest papers and repos is available at https://github.com/LMM101/Awesome-Multimodal-Next-Token-Prediction
format Preprint
id arxiv_https___arxiv_org_abs_2412_18619
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey
Chen, Liang
Wang, Zekun
Ren, Shuhuai
Li, Lei
Zhao, Haozhe
Li, Yunshui
Cai, Zefan
Guo, Hongcheng
Zhang, Lei
Xiong, Yizhe
Zhang, Yichi
Wu, Ruoyu
Dong, Qingxiu
Zhang, Ge
Yang, Jian
Meng, Lingwei
Hu, Shujie
Chen, Yulong
Lin, Junyang
Bai, Shuai
Vlachos, Andreas
Tan, Xu
Zhang, Minjia
Xiao, Wen
Yee, Aaron
Liu, Tianyu
Chang, Baobao
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
Multimedia
Audio and Speech Processing
Building on the foundations of language modeling in natural language processing, Next Token Prediction (NTP) has evolved into a versatile training objective for machine learning tasks across various modalities, achieving considerable success. As Large Language Models (LLMs) have advanced to unify understanding and generation tasks within the textual modality, recent research has shown that tasks from different modalities can also be effectively encapsulated within the NTP framework, transforming the multimodal information into tokens and predict the next one given the context. This survey introduces a comprehensive taxonomy that unifies both understanding and generation within multimodal learning through the lens of NTP. The proposed taxonomy covers five key aspects: Multimodal tokenization, MMNTP model architectures, unified task representation, datasets \& evaluation, and open challenges. This new taxonomy aims to aid researchers in their exploration of multimodal intelligence. An associated GitHub repository collecting the latest papers and repos is available at https://github.com/LMM101/Awesome-Multimodal-Next-Token-Prediction
title Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
Multimedia
Audio and Speech Processing
url https://arxiv.org/abs/2412.18619