Emerging Properties in Unified Multimodal Pretraining
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916866452094976 |
|---|---|
| author | Deng, Chaorui Zhu, Deyao Li, Kunchang Gou, Chenhui Li, Feng Wang, Zeyu Zhong, Shu Yu, Weihao Nie, Xiaonan Song, Ziang Shi, Guang Fan, Haoqi |
| author_facet | Deng, Chaorui Zhu, Deyao Li, Kunchang Gou, Chenhui Li, Feng Wang, Zeyu Zhong, Shu Yu, Weihao Nie, Xiaonan Song, Ziang Shi, Guang Fan, Haoqi |
| contents | Unifying multimodal understanding and generation has shown impressive capabilities in cutting-edge proprietary systems. In this work, we introduce BAGEL, an open-source foundational model that natively supports multimodal understanding and generation. BAGEL is a unified, decoder-only model pretrained on trillions of tokens curated from large-scale interleaved text, image, video, and web data. When scaled with such diverse multimodal interleaved data, BAGEL exhibits emerging capabilities in complex multimodal reasoning. As a result, it significantly outperforms open-source unified models in both multimodal generation and understanding across standard benchmarks, while exhibiting advanced multimodal reasoning abilities such as free-form image manipulation, future frame prediction, 3D manipulation, and world navigation. In the hope of facilitating further opportunities for multimodal research, we share the key findings, pretraining details, data creation protocal, and release our code and checkpoints to the community. The project page is at https://bagel-ai.org/ |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_14683 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Emerging Properties in Unified Multimodal Pretraining Deng, Chaorui Zhu, Deyao Li, Kunchang Gou, Chenhui Li, Feng Wang, Zeyu Zhong, Shu Yu, Weihao Nie, Xiaonan Song, Ziang Shi, Guang Fan, Haoqi Computer Vision and Pattern Recognition Unifying multimodal understanding and generation has shown impressive capabilities in cutting-edge proprietary systems. In this work, we introduce BAGEL, an open-source foundational model that natively supports multimodal understanding and generation. BAGEL is a unified, decoder-only model pretrained on trillions of tokens curated from large-scale interleaved text, image, video, and web data. When scaled with such diverse multimodal interleaved data, BAGEL exhibits emerging capabilities in complex multimodal reasoning. As a result, it significantly outperforms open-source unified models in both multimodal generation and understanding across standard benchmarks, while exhibiting advanced multimodal reasoning abilities such as free-form image manipulation, future frame prediction, 3D manipulation, and world navigation. In the hope of facilitating further opportunities for multimodal research, we share the key findings, pretraining details, data creation protocal, and release our code and checkpoints to the community. The project page is at https://bagel-ai.org/ |
| title | Emerging Properties in Unified Multimodal Pretraining |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2505.14683 |