OmniGen2: Towards Instruction-Aligned Multimodal Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Chenyuan, Zheng, Pengfei, Yan, Ruiran, Xiao, Shitao, Luo, Xin, Wang, Yueze, Li, Wanli, Jiang, Xiyan, Liu, Yexin, Zhou, Junjie, Liu, Ze, Xia, Ziyi, Li, Chaofan, Deng, Haoge, Wang, Jiahao, Luo, Kun, Zhang, Bo, Lian, Defu, Wang, Xinlong, Wang, Zhongyuan, Huang, Tiejun, Liu, Zheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OmniGen: Unified Image Generation
by: Xiao, Shitao, et al.
Published: (2024)
by: Xiao, Shitao, et al.
Published: (2024)
EditScore: Unlocking Online RL for Image Editing via High-Fidelity Reward Modeling
by: Luo, Xin, et al.
Published: (2025)
by: Luo, Xin, et al.
Published: (2025)
OmniGen: Unified Multimodal Sensor Generation for Autonomous Driving
by: Tang, Tao, et al.
Published: (2025)
by: Tang, Tao, et al.
Published: (2025)
O1 Embedder: Let Retrievers Think Before Action
by: Yan, Ruiran, et al.
Published: (2025)
by: Yan, Ruiran, et al.
Published: (2025)
Llama2Vec: Unsupervised Adaptation of Large Language Models for Dense Retrieval
by: Liu, Zheng, et al.
Published: (2023)
by: Liu, Zheng, et al.
Published: (2023)
Lighter And Better: Towards Flexible Context Adaptation For Retrieval Augmented Generation
by: Liu, Zheng, et al.
Published: (2024)
by: Liu, Zheng, et al.
Published: (2024)
MegaPairs: Massive Data Synthesis For Universal Multimodal Retrieval
by: Zhou, Junjie, et al.
Published: (2024)
by: Zhou, Junjie, et al.
Published: (2024)
Matryoshka Re-Ranker: A Flexible Re-Ranking Architecture With Configurable Depth and Width
by: Liu, Zheng, et al.
Published: (2025)
by: Liu, Zheng, et al.
Published: (2025)
You See it, You Got it: Learning 3D Creation on Pose-Free Videos at Scale
by: Ma, Baorui, et al.
Published: (2024)
by: Ma, Baorui, et al.
Published: (2024)
OmniUMI: Towards Physically Grounded Robot Learning via Human-Aligned Multimodal Interaction
by: Luo, Shaqi, et al.
Published: (2026)
by: Luo, Shaqi, et al.
Published: (2026)
LINA: Linear Autoregressive Image Generative Models with Continuous Tokens
by: Wang, Jiahao, et al.
Published: (2026)
by: Wang, Jiahao, et al.
Published: (2026)
Efficient Multimodal Learning from Data-centric Perspective
by: He, Muyang, et al.
Published: (2024)
by: He, Muyang, et al.
Published: (2024)
Making Text Embedders Few-Shot Learners
by: Li, Chaofan, et al.
Published: (2024)
by: Li, Chaofan, et al.
Published: (2024)
AIR-Bench: Automated Heterogeneous Information Retrieval Benchmark
by: Chen, Jianlyu, et al.
Published: (2024)
by: Chen, Jianlyu, et al.
Published: (2024)
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models
by: Diao, Haiwen, et al.
Published: (2025)
by: Diao, Haiwen, et al.
Published: (2025)
Large Language Models as Foundations for Next-Gen Dense Retrieval: A Comprehensive Empirical Assessment
by: Luo, Kun, et al.
Published: (2024)
by: Luo, Kun, et al.
Published: (2024)
M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation
by: Chen, Jianlv, et al.
Published: (2024)
by: Chen, Jianlv, et al.
Published: (2024)
Emu3.5: Native Multimodal Models are World Learners
by: Cui, Yufeng, et al.
Published: (2025)
by: Cui, Yufeng, et al.
Published: (2025)
AlignVid: Training-Free Attention Scaling for Semantic Fidelity in Text-Guided Image-to-Video Generation
by: Liu, Yexin, et al.
Published: (2025)
by: Liu, Yexin, et al.
Published: (2025)
BGE Landmark Embedding: A Chunking-Free Embedding Method For Retrieval Augmented Long-Context Large Language Models
by: Luo, Kun, et al.
Published: (2024)
by: Luo, Kun, et al.
Published: (2024)
Generative Multimodal Models are In-Context Learners
by: Sun, Quan, et al.
Published: (2023)
by: Sun, Quan, et al.
Published: (2023)
Towards A Generalist Code Embedding Model Based On Massive Data Synthesis
by: Li, Chaofan, et al.
Published: (2025)
by: Li, Chaofan, et al.
Published: (2025)
ReasonEmbed: Enhanced Text Embeddings for Reasoning-Intensive Document Retrieval
by: Chen, Jianlyu, et al.
Published: (2025)
by: Chen, Jianlyu, et al.
Published: (2025)
Reinforced Information Retrieval
by: Li, Chaofan, et al.
Published: (2025)
by: Li, Chaofan, et al.
Published: (2025)
MR$^2$-Bench: Going Beyond Matching to Reasoning in Multimodal Retrieval
by: Zhou, Junjie, et al.
Published: (2025)
by: Zhou, Junjie, et al.
Published: (2025)
DeepXiv-SDK: An Agentic Data Interface for Scientific Literature
by: Qian, Hongjin, et al.
Published: (2026)
by: Qian, Hongjin, et al.
Published: (2026)
Emu: Generative Pretraining in Multimodality
by: Sun, Quan, et al.
Published: (2023)
by: Sun, Quan, et al.
Published: (2023)
Unveiling Encoder-Free Vision-Language Models
by: Diao, Haiwen, et al.
Published: (2024)
by: Diao, Haiwen, et al.
Published: (2024)
MYCloth: Towards Intelligent and Interactive Online T-Shirt Customization based on User's Preference
by: Liu, Yexin, et al.
Published: (2024)
by: Liu, Yexin, et al.
Published: (2024)
UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models
by: Wang, Jiaqi, et al.
Published: (2026)
by: Wang, Jiaqi, et al.
Published: (2026)
GoodSAM++: Bridging Domain and Capacity Gaps via Segment Anything Model for Panoramic Semantic Segmentation
by: Zhang, Weiming, et al.
Published: (2024)
by: Zhang, Weiming, et al.
Published: (2024)
GoodSAM: Bridging Domain and Capacity Gaps via Segment Anything Model for Distortion-aware Panoramic Semantic Segmentation
by: Zhang, Weiming, et al.
Published: (2024)
by: Zhang, Weiming, et al.
Published: (2024)
Thor: Towards Human-Level Whole-Body Reactions for Intense Contact-Rich Environments
by: Li, Gangyang, et al.
Published: (2025)
by: Li, Gangyang, et al.
Published: (2025)
Retro*: Optimizing LLMs for Reasoning-Intensive Document Retrieval
by: Lan, Junwei, et al.
Published: (2025)
by: Lan, Junwei, et al.
Published: (2025)
Unveiling the Ignorance of MLLMs: Seeing Clearly, Answering Incorrectly
by: Liu, Yexin, et al.
Published: (2024)
by: Liu, Yexin, et al.
Published: (2024)
DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception
by: Li, Xiaotong, et al.
Published: (2024)
by: Li, Xiaotong, et al.
Published: (2024)
InfoFlow: Reinforcing Search Agent Via Reward Density Optimization
by: Luo, Kun, et al.
Published: (2025)
by: Luo, Kun, et al.
Published: (2025)
Any Information Is Just Worth One Single Screenshot: Unifying Search With Visualized Information Retrieval
by: Liu, Ze, et al.
Published: (2025)
by: Liu, Ze, et al.
Published: (2025)
Are the Values of LLMs Structurally Aligned with Humans? A Causal Perspective
by: Kang, Yipeng, et al.
Published: (2024)
by: Kang, Yipeng, et al.
Published: (2024)
A longitudinal dyadic analysis of gender ideology during the transition into parenthood
by: Yexin Zheng, et al.
Published: (2025)
by: Yexin Zheng, et al.
Published: (2025)
Similar Items
-
OmniGen: Unified Image Generation
by: Xiao, Shitao, et al.
Published: (2024) -
EditScore: Unlocking Online RL for Image Editing via High-Fidelity Reward Modeling
by: Luo, Xin, et al.
Published: (2025) -
OmniGen: Unified Multimodal Sensor Generation for Autonomous Driving
by: Tang, Tao, et al.
Published: (2025) -
O1 Embedder: Let Retrievers Think Before Action
by: Yan, Ruiran, et al.
Published: (2025) -
Llama2Vec: Unsupervised Adaptation of Large Language Models for Dense Retrieval
by: Liu, Zheng, et al.
Published: (2023)