BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Chen, Jiuhai, Xu, Zhiyang, Pan, Xichen, Hu, Yushi, Qin, Can, Goldstein, Tom, Huang, Lifu, Zhou, Tianyi, Xie, Saining, Savarese, Silvio, Xue, Le, Xiong, Caiming, Xu, Ran |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
BLIP3o-NEXT: Next Frontier of Native Image Generation
par: Chen, Jiuhai, et autres
Publié: (2025)
par: Chen, Jiuhai, et autres
Publié: (2025)
xGen-MM (BLIP-3): A Family of Open Large Multimodal Models
par: Xue, Le, et autres
Publié: (2024)
par: Xue, Le, et autres
Publié: (2024)
X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning
par: Panagopoulou, Artemis, et autres
Publié: (2023)
par: Panagopoulou, Artemis, et autres
Publié: (2023)
xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs
par: Ryoo, Michael S., et autres
Publié: (2024)
par: Ryoo, Michael S., et autres
Publié: (2024)
BLIP3-KALE: Knowledge Augmented Large-Scale Dense Captions
par: Awadalla, Anas, et autres
Publié: (2024)
par: Awadalla, Anas, et autres
Publié: (2024)
MINT-1T: Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
par: Awadalla, Anas, et autres
Publié: (2024)
par: Awadalla, Anas, et autres
Publié: (2024)
DyMU: Dynamic Merging and Virtual Unmerging for Efficient VLMs
par: Wang, Zhenhailong, et autres
Publié: (2025)
par: Wang, Zhenhailong, et autres
Publié: (2025)
UNIDOC-BENCH: A Unified Benchmark for Document-Centric Multimodal RAG
par: Peng, Xiangyu, et autres
Publié: (2025)
par: Peng, Xiangyu, et autres
Publié: (2025)
Unified Training of Universal Time Series Forecasting Transformers
par: Woo, Gerald, et autres
Publié: (2024)
par: Woo, Gerald, et autres
Publié: (2024)
Enabling High Data Throughput Reinforcement Learning on GPUs: A Domain Agnostic Framework for Data-Driven Scientific Research
par: Lan, Tian, et autres
Publié: (2024)
par: Lan, Tian, et autres
Publié: (2024)
MULTISCRIPT: Multimodal Script Learning for Supporting Open Domain Everyday Tasks
par: Qi, Jingyuan, et autres
Publié: (2023)
par: Qi, Jingyuan, et autres
Publié: (2023)
Shared Imagination: LLMs Hallucinate Alike
par: Zhou, Yilun, et autres
Publié: (2024)
par: Zhou, Yilun, et autres
Publié: (2024)
BOLT: Bootstrap Long Chain-of-Thought in Language Models without Distillation
par: Pang, Bo, et autres
Publié: (2025)
par: Pang, Bo, et autres
Publié: (2025)
CodeXEmbed: A Generalist Embedding Model Family for Multiligual and Multi-task Code Retrieval
par: Liu, Ye, et autres
Publié: (2024)
par: Liu, Ye, et autres
Publié: (2024)
Reasoning Curriculum: Bootstrapping Broad LLM Reasoning from Math
par: Pang, Bo, et autres
Publié: (2025)
par: Pang, Bo, et autres
Publié: (2025)
INDICT: Code Generation with Internal Dialogues of Critiques for Both Security and Helpfulness
par: Le, Hung, et autres
Publié: (2024)
par: Le, Hung, et autres
Publié: (2024)
DialogStudio: Towards Richest and Most Diverse Unified Dataset Collection for Conversational AI
par: Zhang, Jianguo, et autres
Publié: (2023)
par: Zhang, Jianguo, et autres
Publié: (2023)
ULIP-2: Towards Scalable Multimodal Pre-training for 3D Understanding
par: Xue, Le, et autres
Publié: (2023)
par: Xue, Le, et autres
Publié: (2023)
Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation
par: Xu, Zhiyang, et autres
Publié: (2025)
par: Xu, Zhiyang, et autres
Publié: (2025)
Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
par: Tong, Shengbang, et autres
Publié: (2024)
par: Tong, Shengbang, et autres
Publié: (2024)
Transfer between Modalities with MetaQueries
par: Pan, Xichen, et autres
Publié: (2025)
par: Pan, Xichen, et autres
Publié: (2025)
UniHGKR: Unified Instruction-aware Heterogeneous Knowledge Retrievers
par: Min, Dehai, et autres
Publié: (2024)
par: Min, Dehai, et autres
Publié: (2024)
Text2Data: Low-Resource Data Generation with Textual Control
par: Wang, Shiyu, et autres
Publié: (2024)
par: Wang, Shiyu, et autres
Publié: (2024)
Future Optical Flow Prediction Improves Robot Control & Video Generation
par: Ranasinghe, Kanchana, et autres
Publié: (2026)
par: Ranasinghe, Kanchana, et autres
Publié: (2026)
Hierarchical Point Attention for Indoor 3D Object Detection
par: Shu, Manli, et autres
Publié: (2023)
par: Shu, Manli, et autres
Publié: (2023)
CodeTree: Agent-guided Tree Search for Code Generation with Large Language Models
par: Li, Jierui, et autres
Publié: (2024)
par: Li, Jierui, et autres
Publié: (2024)
ProVision: Programmatically Scaling Vision-centric Instruction Data for Multimodal Language Models
par: Zhang, Jieyu, et autres
Publié: (2024)
par: Zhang, Jieyu, et autres
Publié: (2024)
GenQA: Generating Millions of Instructions from a Handful of Prompts
par: Chen, Jiuhai, et autres
Publié: (2024)
par: Chen, Jiuhai, et autres
Publié: (2024)
LaTtE-Flow: Layerwise Timestep-Expert Flow-based Transformer
par: Shen, Ying, et autres
Publié: (2025)
par: Shen, Ying, et autres
Publié: (2025)
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents
par: Nguyen, Xuan-Phi, et autres
Publié: (2025)
par: Nguyen, Xuan-Phi, et autres
Publié: (2025)
GIFT-Eval: A Benchmark For General Time Series Forecasting Model Evaluation
par: Aksu, Taha, et autres
Publié: (2024)
par: Aksu, Taha, et autres
Publié: (2024)
Reward-Guided Speculative Decoding for Efficient LLM Reasoning
par: Liao, Baohao, et autres
Publié: (2025)
par: Liao, Baohao, et autres
Publié: (2025)
Entropy-Based Block Pruning for Efficient Large Language Models
par: Yang, Liangwei, et autres
Publié: (2025)
par: Yang, Liangwei, et autres
Publié: (2025)
PerfCodeGen: Improving Performance of LLM Generated Code with Execution Feedback
par: Peng, Yun, et autres
Publié: (2024)
par: Peng, Yun, et autres
Publié: (2024)
Asynchronous Tool Usage for Real-Time Agents
par: Ginart, Antonio A., et autres
Publié: (2024)
par: Ginart, Antonio A., et autres
Publié: (2024)
HIVE: Harnessing Human Feedback for Instructional Visual Editing
par: Zhang, Shu, et autres
Publié: (2023)
par: Zhang, Shu, et autres
Publié: (2023)
Multimodal Instruction Tuning with Conditional Mixture of LoRA
par: Shen, Ying, et autres
Publié: (2024)
par: Shen, Ying, et autres
Publié: (2024)
LZ Penalty: An information-theoretic repetition penalty for autoregressive language models
par: Ginart, Antonio A., et autres
Publié: (2025)
par: Ginart, Antonio A., et autres
Publié: (2025)
Multi-Objective Linguistic Control of Large Language Models
par: Nguyen, Dang, et autres
Publié: (2024)
par: Nguyen, Dang, et autres
Publié: (2024)
AR-RAG: Autoregressive Retrieval Augmentation for Image Generation
par: Qi, Jingyuan, et autres
Publié: (2025)
par: Qi, Jingyuan, et autres
Publié: (2025)
Documents similaires
-
BLIP3o-NEXT: Next Frontier of Native Image Generation
par: Chen, Jiuhai, et autres
Publié: (2025) -
xGen-MM (BLIP-3): A Family of Open Large Multimodal Models
par: Xue, Le, et autres
Publié: (2024) -
X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning
par: Panagopoulou, Artemis, et autres
Publié: (2023) -
xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs
par: Ryoo, Michael S., et autres
Publié: (2024) -
BLIP3-KALE: Knowledge Augmented Large-Scale Dense Captions
par: Awadalla, Anas, et autres
Publié: (2024)