PIN: A Knowledge-Intensive Dataset for Paired and Interleaved Multimodal Documents

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Junjie, Zhang, Yuxiang, Liu, Minghao, Zhang, Yin, Ji, Yatai, Xuan, Weihao, Lin, Nie, Zhu, Kang, Lin, Zhiqiang, Ren, Yiming, Jiang, Chunyang, Yu, Yiyao, Wang, Zekun, Wang, Tiezhen, Huang, Wenhao, Fu, Jie, Lin, Qunshu, Yang, Yujiu, Zhang, Ge, Yuan, Ruibin, Chen, Bei, Chen, Wenhu
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918137748783104
author Wang, Junjie
Zhang, Yuxiang
Liu, Minghao
Zhang, Yin
Ji, Yatai
Xuan, Weihao
Lin, Nie
Zhu, Kang
Lin, Zhiqiang
Ren, Yiming
Jiang, Chunyang
Yu, Yiyao
Wang, Zekun
Wang, Tiezhen
Huang, Wenhao
Fu, Jie
Lin, Qunshu
Yang, Yujiu
Zhang, Ge
Yuan, Ruibin
Chen, Bei
Chen, Wenhu
author_facet Wang, Junjie
Zhang, Yuxiang
Liu, Minghao
Zhang, Yin
Ji, Yatai
Xuan, Weihao
Lin, Nie
Zhu, Kang
Lin, Zhiqiang
Ren, Yiming
Jiang, Chunyang
Yu, Yiyao
Wang, Zekun
Wang, Tiezhen
Huang, Wenhao
Fu, Jie
Lin, Qunshu
Yang, Yujiu
Zhang, Ge
Yuan, Ruibin
Chen, Bei
Chen, Wenhu
contents Recent advancements in large multimodal models (LMMs) have leveraged extensive multimodal datasets to enhance capabilities in complex knowledge-driven tasks. However, persistent challenges in perceptual and reasoning errors limit their efficacy, particularly in interpreting intricate visual data and deducing multimodal relationships. To address these issues, we introduce PIN (Paired and INterleaved multimodal documents), a novel data format designed to foster a deeper integration of visual and textual knowledge. The PIN format uniquely combines semantically rich Markdown files, which preserve fine-grained textual structures, with holistic overall images that capture the complete document layout. Following this format, we construct and release two large-scale, open-source datasets: PIN-200M (~200 million documents) and PIN-14M (~14 million), compiled from diverse web and scientific sources in both English and Chinese. To maximize usability, we provide detailed statistical analyses and equip the datasets with quality signals, enabling researchers to easily filter and select data for specific tasks. Our work provides the community with a versatile data format and substantial resources, offering a foundation for new research in pre-training strategies and the development of more powerful knowledge-intensive LMMs.
format Preprint
id arxiv_https___arxiv_org_abs_2406_13923
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle PIN: A Knowledge-Intensive Dataset for Paired and Interleaved Multimodal Documents
Wang, Junjie
Zhang, Yuxiang
Liu, Minghao
Zhang, Yin
Ji, Yatai
Xuan, Weihao
Lin, Nie
Zhu, Kang
Lin, Zhiqiang
Ren, Yiming
Jiang, Chunyang
Yu, Yiyao
Wang, Zekun
Wang, Tiezhen
Huang, Wenhao
Fu, Jie
Lin, Qunshu
Yang, Yujiu
Zhang, Ge
Yuan, Ruibin
Chen, Bei
Chen, Wenhu
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Multimedia
Recent advancements in large multimodal models (LMMs) have leveraged extensive multimodal datasets to enhance capabilities in complex knowledge-driven tasks. However, persistent challenges in perceptual and reasoning errors limit their efficacy, particularly in interpreting intricate visual data and deducing multimodal relationships. To address these issues, we introduce PIN (Paired and INterleaved multimodal documents), a novel data format designed to foster a deeper integration of visual and textual knowledge. The PIN format uniquely combines semantically rich Markdown files, which preserve fine-grained textual structures, with holistic overall images that capture the complete document layout. Following this format, we construct and release two large-scale, open-source datasets: PIN-200M (~200 million documents) and PIN-14M (~14 million), compiled from diverse web and scientific sources in both English and Chinese. To maximize usability, we provide detailed statistical analyses and equip the datasets with quality signals, enabling researchers to easily filter and select data for specific tasks. Our work provides the community with a versatile data format and substantial resources, offering a foundation for new research in pre-training strategies and the development of more powerful knowledge-intensive LMMs.
title PIN: A Knowledge-Intensive Dataset for Paired and Interleaved Multimodal Documents
topic Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2406.13923