OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Qingyun, Chen, Zhe, Wang, Weiyun, Wang, Wenhai, Ye, Shenglong, Jin, Zhenjiang, Chen, Guanzhou, He, Yinan, Gao, Zhangwei, Cui, Erfei, Yu, Jiashuo, Tian, Hao, Zhou, Jiasheng, Xu, Chao, Wang, Bin, Wei, Xingjian, Li, Wei, Zhang, Wenjian, Zhang, Bo, Cai, Pinlong, Wen, Licheng, Yan, Xiangchao, Li, Zhenxiang, Chu, Pei, Wang, Yi, Dou, Min, Tian, Changyao, Zhu, Xizhou, Lu, Lewei, Chen, Yushi, He, Junjun, Tu, Zhongying, Lu, Tong, Wang, Yali, Wang, Limin, Lin, Dahua, Qiao, Yu, Shi, Botian, He, Conghui, Dai, Jifeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance
von: Gao, Zhangwei, et al.
Veröffentlicht: (2024)
von: Gao, Zhangwei, et al.
Veröffentlicht: (2024)
MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer
von: Tian, Changyao, et al.
Veröffentlicht: (2024)
von: Tian, Changyao, et al.
Veröffentlicht: (2024)
MetaCaptioner: Towards Generalist Visual Captioning with Open-source Suites
von: Lei, Zhenxin, et al.
Veröffentlicht: (2025)
von: Lei, Zhenxin, et al.
Veröffentlicht: (2025)
MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity
von: Liu, Yangzhou, et al.
Veröffentlicht: (2024)
von: Liu, Yangzhou, et al.
Veröffentlicht: (2024)
Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization
von: Wang, Weiyun, et al.
Veröffentlicht: (2024)
von: Wang, Weiyun, et al.
Veröffentlicht: (2024)
Docopilot: Improving Multimodal Models for Document-Level Understanding
von: Duan, Yuchen, et al.
Veröffentlicht: (2025)
von: Duan, Yuchen, et al.
Veröffentlicht: (2025)
Vision-RWKV: Efficient and Scalable Visual Perception with RWKV-Like Architectures
von: Duan, Yuchen, et al.
Veröffentlicht: (2024)
von: Duan, Yuchen, et al.
Veröffentlicht: (2024)
How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
von: Chen, Zhe, et al.
Veröffentlicht: (2024)
von: Chen, Zhe, et al.
Veröffentlicht: (2024)
VisualPRM: An Effective Process Reward Model for Multimodal Reasoning
von: Wang, Weiyun, et al.
Veröffentlicht: (2025)
von: Wang, Weiyun, et al.
Veröffentlicht: (2025)
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models
von: Luo, Gen, et al.
Veröffentlicht: (2025)
von: Luo, Gen, et al.
Veröffentlicht: (2025)
The All-Seeing Project V2: Towards General Relation Comprehension of the Open World
von: Wang, Weiyun, et al.
Veröffentlicht: (2024)
von: Wang, Weiyun, et al.
Veröffentlicht: (2024)
MathPile: A Billion-Token-Scale Pretraining Corpus for Math
von: Wang, Zengzhi, et al.
Veröffentlicht: (2023)
von: Wang, Zengzhi, et al.
Veröffentlicht: (2023)
ADDP: Learning General Representations for Image Recognition and Generation with Alternating Denoising Diffusion Process
von: Tian, Changyao, et al.
Veröffentlicht: (2023)
von: Tian, Changyao, et al.
Veröffentlicht: (2023)
DriveMLM: Aligning Multi-Modal Large Language Models with Behavioral Planning States for Autonomous Driving
von: Cui, Erfei, et al.
Veröffentlicht: (2023)
von: Cui, Erfei, et al.
Veröffentlicht: (2023)
VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models
von: Xu, Weiye, et al.
Veröffentlicht: (2025)
von: Xu, Weiye, et al.
Veröffentlicht: (2025)
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces
von: Luo, Gen, et al.
Veröffentlicht: (2025)
von: Luo, Gen, et al.
Veröffentlicht: (2025)
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding
von: Li, Hao, et al.
Veröffentlicht: (2024)
von: Li, Hao, et al.
Veröffentlicht: (2024)
PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models
von: Yang, Chenyu, et al.
Veröffentlicht: (2024)
von: Yang, Chenyu, et al.
Veröffentlicht: (2024)
ScaleEdit-12M: Scaling Open-Source Image Editing Data Generation via Multi-Agent Framework
von: Chen, Guanzhou, et al.
Veröffentlicht: (2026)
von: Chen, Guanzhou, et al.
Veröffentlicht: (2026)
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
von: Chen, Zhe, et al.
Veröffentlicht: (2024)
von: Chen, Zhe, et al.
Veröffentlicht: (2024)
Demystify Transformers & Convolutions in Modern Image Deep Networks
von: Hu, Xiaowei, et al.
Veröffentlicht: (2022)
von: Hu, Xiaowei, et al.
Veröffentlicht: (2022)
NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints
von: Tian, Changyao, et al.
Veröffentlicht: (2025)
von: Tian, Changyao, et al.
Veröffentlicht: (2025)
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding
von: Wang, Zhaokai, et al.
Veröffentlicht: (2025)
von: Wang, Zhaokai, et al.
Veröffentlicht: (2025)
Efficient Deformable ConvNets: Rethinking Dynamic and Sparse Operator for Vision Applications
von: Xiong, Yuwen, et al.
Veröffentlicht: (2024)
von: Xiong, Yuwen, et al.
Veröffentlicht: (2024)
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
von: Zhu, Jinguo, et al.
Veröffentlicht: (2025)
von: Zhu, Jinguo, et al.
Veröffentlicht: (2025)
Rethinking Multi-view Representation Learning via Distilled Disentangling
von: Ke, Guanzhou, et al.
Veröffentlicht: (2024)
von: Ke, Guanzhou, et al.
Veröffentlicht: (2024)
GFPL: Generative Federated Prototype Learning for Resource-Constrained and Data-Imbalanced Vision Task
von: Lu, Shiwei, et al.
Veröffentlicht: (2026)
von: Lu, Shiwei, et al.
Veröffentlicht: (2026)
Sequential Diffusion Language Models
von: Liu, Yangzhou, et al.
Veröffentlicht: (2025)
von: Liu, Yangzhou, et al.
Veröffentlicht: (2025)
Needle In A Multimodal Haystack
von: Wang, Weiyun, et al.
Veröffentlicht: (2024)
von: Wang, Weiyun, et al.
Veröffentlicht: (2024)
Learning 1D Causal Visual Representation with De-focus Attention Networks
von: Tao, Chenxin, et al.
Veröffentlicht: (2024)
von: Tao, Chenxin, et al.
Veröffentlicht: (2024)
Characterizing and Mitigating False-Positive Bug Reports in the Linux Kernel
von: Tian, Jiashuo, et al.
Veröffentlicht: (2026)
von: Tian, Jiashuo, et al.
Veröffentlicht: (2026)
TACOMORE: Leveraging the Potential of LLMs in Corpus-based Discourse Analysis with Prompt Engineering
von: Li, Bingru, et al.
Veröffentlicht: (2024)
von: Li, Bingru, et al.
Veröffentlicht: (2024)
GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks
von: Zhang, Yu, et al.
Veröffentlicht: (2024)
von: Zhang, Yu, et al.
Veröffentlicht: (2024)
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos
von: Yu, Jiashuo, et al.
Veröffentlicht: (2025)
von: Yu, Jiashuo, et al.
Veröffentlicht: (2025)
GenExam: A Multidisciplinary Text-to-Image Exam
von: Wang, Zhaokai, et al.
Veröffentlicht: (2025)
von: Wang, Zhaokai, et al.
Veröffentlicht: (2025)
Vision Model Pre-training on Interleaved Image-Text Data via Latent Compression Learning
von: Yang, Chenyu, et al.
Veröffentlicht: (2024)
von: Yang, Chenyu, et al.
Veröffentlicht: (2024)
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding
von: Tao, Chenxin, et al.
Veröffentlicht: (2024)
von: Tao, Chenxin, et al.
Veröffentlicht: (2024)
FALCON: Feedback-driven Adaptive Long/short-term memory reinforced Coding Optimization system
von: Li, Zeyuan, et al.
Veröffentlicht: (2024)
von: Li, Zeyuan, et al.
Veröffentlicht: (2024)
MSR-86K: An Evolving, Multilingual Corpus with 86,300 Hours of Transcribed Audio for Speech Recognition Research
von: Li, Song, et al.
Veröffentlicht: (2024)
von: Li, Song, et al.
Veröffentlicht: (2024)
Auto MC-Reward: Automated Dense Reward Design with Large Language Models for Minecraft
von: Li, Hao, et al.
Veröffentlicht: (2023)
von: Li, Hao, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance
von: Gao, Zhangwei, et al.
Veröffentlicht: (2024) -
MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer
von: Tian, Changyao, et al.
Veröffentlicht: (2024) -
MetaCaptioner: Towards Generalist Visual Captioning with Open-source Suites
von: Lei, Zhenxin, et al.
Veröffentlicht: (2025) -
MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity
von: Liu, Yangzhou, et al.
Veröffentlicht: (2024) -
Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization
von: Wang, Weiyun, et al.
Veröffentlicht: (2024)