MINT-1T: Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
Fuente:
arXiv
Saved in:
| Main Authors: | Awadalla, Anas, Xue, Le, Lo, Oscar, Shu, Manli, Lee, Hannah, Guha, Etash Kumar, Jordan, Matt, Shen, Sheng, Awadalla, Mohamed, Savarese, Silvio, Xiong, Caiming, Xu, Ran, Choi, Yejin, Schmidt, Ludwig |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BLIP3-KALE: Knowledge Augmented Large-Scale Dense Captions
by: Awadalla, Anas, et al.
Published: (2024)
by: Awadalla, Anas, et al.
Published: (2024)
Certainly Uncertain: A Benchmark and Metric for Multimodal Epistemic and Aleatoric Awareness
by: Chandu, Khyathi Raghavi, et al.
Published: (2024)
by: Chandu, Khyathi Raghavi, et al.
Published: (2024)
xGen-MM (BLIP-3): A Family of Open Large Multimodal Models
by: Xue, Le, et al.
Published: (2024)
by: Xue, Le, et al.
Published: (2024)
xGen-VideoSyn-1: High-fidelity Text-to-Video Synthesis with Compressed Representations
by: Qin, Can, et al.
Published: (2024)
by: Qin, Can, et al.
Published: (2024)
Flexibility Characterization of Sustainable Power Systems in Demand Space: A Data-Driven Inverse Optimization Approach
by: Awadalla, Mohamed, et al.
Published: (2022)
by: Awadalla, Mohamed, et al.
Published: (2022)
AFSR: An Adaptive Factor Score Regression Framework for High‐Dimensional Correlation Matrix Estimation
by: Muath Awadalla, et al.
Published: (2026)
by: Muath Awadalla, et al.
Published: (2026)
xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs
by: Ryoo, Michael S., et al.
Published: (2024)
by: Ryoo, Michael S., et al.
Published: (2024)
Solving Robust MDPs through No-Regret Dynamics
by: Guha, Etash Kumar
Published: (2023)
by: Guha, Etash Kumar
Published: (2023)
Infini-gram: Scaling Unbounded n-gram Language Models to a Trillion Tokens
by: Liu, Jiacheng, et al.
Published: (2024)
by: Liu, Jiacheng, et al.
Published: (2024)
On the Diminishing Returns of Width for Continual Learning
by: Guha, Etash, et al.
Published: (2024)
by: Guha, Etash, et al.
Published: (2024)
ProVision: Programmatically Scaling Vision-centric Instruction Data for Multimodal Language Models
by: Zhang, Jieyu, et al.
Published: (2024)
by: Zhang, Jieyu, et al.
Published: (2024)
Enabling High Data Throughput Reinforcement Learning on GPUs: A Domain Agnostic Framework for Data-Driven Scientific Research
by: Lan, Tian, et al.
Published: (2024)
by: Lan, Tian, et al.
Published: (2024)
Existence and Stability of Ulam–Hyers and Generalized Ulam–Hyers for the Generalized Langevin–Sturm–Liouville Equation Involving Generalized Liouville–Caputo Type
by: Muthaiah Subramanian, et al.
Published: (2025)
by: Muthaiah Subramanian, et al.
Published: (2025)
DyMU: Dynamic Merging and Virtual Unmerging for Efficient VLMs
by: Wang, Zhenhailong, et al.
Published: (2025)
by: Wang, Zhenhailong, et al.
Published: (2025)
Shared Imagination: LLMs Hallucinate Alike
by: Zhou, Yilun, et al.
Published: (2024)
by: Zhou, Yilun, et al.
Published: (2024)
MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping
by: Shan, Xiaojun, et al.
Published: (2025)
by: Shan, Xiaojun, et al.
Published: (2025)
xGen-small Technical Report
by: Nijkamp, Erik, et al.
Published: (2025)
by: Nijkamp, Erik, et al.
Published: (2025)
Reasoning Curriculum: Bootstrapping Broad LLM Reasoning from Math
by: Pang, Bo, et al.
Published: (2025)
by: Pang, Bo, et al.
Published: (2025)
INDICT: Code Generation with Internal Dialogues of Critiques for Both Security and Helpfulness
by: Le, Hung, et al.
Published: (2024)
by: Le, Hung, et al.
Published: (2024)
NitroGen: An Open Foundation Model for Generalist Gaming Agents
by: Magne, Loïc, et al.
Published: (2026)
by: Magne, Loïc, et al.
Published: (2026)
ULIP-2: Towards Scalable Multimodal Pre-training for 3D Understanding
by: Xue, Le, et al.
Published: (2023)
by: Xue, Le, et al.
Published: (2023)
Existence Results for the System of Fractional-Order Sequential Integrodifferential Equations via Liouville–Caputo Sense
by: Muath Awadalla, et al.
Published: (2024)
by: Muath Awadalla, et al.
Published: (2024)
Mind the Gap No More: Achieving Zero-Gap Multimodal Integration via One Tokenizer
by: Li, Yanan, et al.
Published: (2026)
by: Li, Yanan, et al.
Published: (2026)
BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset
by: Chen, Jiuhai, et al.
Published: (2025)
by: Chen, Jiuhai, et al.
Published: (2025)
BOLT: Bootstrap Long Chain-of-Thought in Language Models without Distillation
by: Pang, Bo, et al.
Published: (2025)
by: Pang, Bo, et al.
Published: (2025)
Unified Training of Universal Time Series Forecasting Transformers
by: Woo, Gerald, et al.
Published: (2024)
by: Woo, Gerald, et al.
Published: (2024)
CodeTree: Agent-guided Tree Search for Code Generation with Large Language Models
by: Li, Jierui, et al.
Published: (2024)
by: Li, Jierui, et al.
Published: (2024)
Efficiently Editing Mixture-of-Experts Models with Compressed Experts
by: He, Yifei, et al.
Published: (2025)
by: He, Yifei, et al.
Published: (2025)
Professional Governance and the Limits of Recognition
by: Mohammed Mustafa, et al.
Published: (2026)
by: Mohammed Mustafa, et al.
Published: (2026)
olmOCR: Unlocking Trillions of Tokens in PDFs with Vision Language Models
by: Poznanski, Jake, et al.
Published: (2025)
by: Poznanski, Jake, et al.
Published: (2025)
PerfCodeGen: Improving Performance of LLM Generated Code with Execution Feedback
by: Peng, Yun, et al.
Published: (2024)
by: Peng, Yun, et al.
Published: (2024)
Asynchronous Tool Usage for Real-Time Agents
by: Ginart, Antonio A., et al.
Published: (2024)
by: Ginart, Antonio A., et al.
Published: (2024)
Are aligned neural networks adversarially aligned?
by: Carlini, Nicholas, et al.
Published: (2023)
by: Carlini, Nicholas, et al.
Published: (2023)
A Paradigm Shift in Machine Translation: Boosting Translation Performance of Large Language Models
by: Xu, Haoran, et al.
Published: (2023)
by: Xu, Haoran, et al.
Published: (2023)
Intern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale
by: Zou, Yicheng, et al.
Published: (2026)
by: Zou, Yicheng, et al.
Published: (2026)
Exploring and Applying Audio-Based Sentiment Analysis in Music
by: Jhanji, Etash
Published: (2024)
by: Jhanji, Etash
Published: (2024)
UNIDOC-BENCH: A Unified Benchmark for Document-Centric Multimodal RAG
by: Peng, Xiangyu, et al.
Published: (2025)
by: Peng, Xiangyu, et al.
Published: (2025)
OLMoTrace: Tracing Language Model Outputs Back to Trillions of Training Tokens
by: Liu, Jiacheng, et al.
Published: (2025)
by: Liu, Jiacheng, et al.
Published: (2025)
MINT: Multimodal Imaging-to-Speech Knowledge Transfer for Early Alzheimer's Screening
by: Ahire, Vrushank, et al.
Published: (2026)
by: Ahire, Vrushank, et al.
Published: (2026)
LZ Penalty: An information-theoretic repetition penalty for autoregressive language models
by: Ginart, Antonio A., et al.
Published: (2025)
by: Ginart, Antonio A., et al.
Published: (2025)
Similar Items
-
BLIP3-KALE: Knowledge Augmented Large-Scale Dense Captions
by: Awadalla, Anas, et al.
Published: (2024) -
Certainly Uncertain: A Benchmark and Metric for Multimodal Epistemic and Aleatoric Awareness
by: Chandu, Khyathi Raghavi, et al.
Published: (2024) -
xGen-MM (BLIP-3): A Family of Open Large Multimodal Models
by: Xue, Le, et al.
Published: (2024) -
xGen-VideoSyn-1: High-fidelity Text-to-Video Synthesis with Compressed Representations
by: Qin, Can, et al.
Published: (2024) -
Flexibility Characterization of Sustainable Power Systems in Demand Space: A Data-Driven Inverse Optimization Approach
by: Awadalla, Mohamed, et al.
Published: (2022)