Towards Unified Multimodal Interleaved Generation via Group Relative Policy Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Nie, Ming, Wang, Chunwei, Han, Jianhua, Xu, Hang, Zhang, Li |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reason2Drive: Towards Interpretable and Chain-based Reasoning for Autonomous Driving
by: Nie, Ming, et al.
Published: (2023)
by: Nie, Ming, et al.
Published: (2023)
KFFocus: Highlighting Keyframes for Enhanced Video Understanding
by: Nie, Ming, et al.
Published: (2025)
by: Nie, Ming, et al.
Published: (2025)
SlowFocus: Enhancing Fine-grained Temporal Understanding in Video LLM
by: Nie, Ming, et al.
Published: (2026)
by: Nie, Ming, et al.
Published: (2026)
SemHiTok: A Unified Image Tokenizer via Semantic-Guided Hierarchical Codebook for Multimodal Understanding and Generation
by: Chen, Zisheng, et al.
Published: (2025)
by: Chen, Zisheng, et al.
Published: (2025)
UNIT: Unifying Image and Text Recognition in One Vision Encoder
by: Zhu, Yi, et al.
Published: (2024)
by: Zhu, Yi, et al.
Published: (2024)
DuoGen: Towards General Purpose Interleaved Multimodal Generation
by: Shi, Min, et al.
Published: (2026)
by: Shi, Min, et al.
Published: (2026)
InterCoG: Towards Spatially Precise Image Editing with Interleaved Chain-of-Grounding Reasoning
by: Wan, Yecong, et al.
Published: (2026)
by: Wan, Yecong, et al.
Published: (2026)
Brick-Diffusion: Generating Long Videos with Brick-to-Wall Denoising
by: Yuan, Yunlong, et al.
Published: (2025)
by: Yuan, Yunlong, et al.
Published: (2025)
Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation
by: Zhang, Yabo, et al.
Published: (2026)
by: Zhang, Yabo, et al.
Published: (2026)
Group Relative Policy Optimization for Image Captioning
by: Liang, Xu
Published: (2025)
by: Liang, Xu
Published: (2025)
ILLUME+: Illuminating Unified MLLM with Dual Visual Tokenization and Diffusion Refinement
by: Huang, Runhui, et al.
Published: (2025)
by: Huang, Runhui, et al.
Published: (2025)
UniM: A Unified Any-to-Any Interleaved Multimodal Benchmark
by: Li, Yanlin, et al.
Published: (2026)
by: Li, Yanlin, et al.
Published: (2026)
ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance
by: Wang, Chunwei, et al.
Published: (2024)
by: Wang, Chunwei, et al.
Published: (2024)
Breaking Dual Bottlenecks: Evolving Unified Multimodal Models into Self-Adaptive Interleaved Visual Reasoners
by: Liu, Qingyang, et al.
Published: (2026)
by: Liu, Qingyang, et al.
Published: (2026)
Robust Object Detection for Autonomous Driving via Curriculum-Guided Group Relative Policy Optimization
by: Jia, Xu
Published: (2025)
by: Jia, Xu
Published: (2025)
DetCLIPv3: Towards Versatile Generative Open-vocabulary Object Detection
by: Yao, Lewei, et al.
Published: (2024)
by: Yao, Lewei, et al.
Published: (2024)
UniGRPO: Unified Policy Optimization for Reasoning-Driven Visual Generation
by: Liu, Jie, et al.
Published: (2026)
by: Liu, Jie, et al.
Published: (2026)
ARMOR: Empowering Multimodal Understanding Model with Interleaved Multimodal Generation Capability
by: Sun, Jianwen, et al.
Published: (2025)
by: Sun, Jianwen, et al.
Published: (2025)
OmniMoGen: Unifying Human Motion Generation via Learning from Interleaved Text-Motion Instructions
by: Bu, Wendong, et al.
Published: (2025)
by: Bu, Wendong, et al.
Published: (2025)
Towards Generalized Multi-Image Editing for Unified Multimodal Models
by: Xu, Pengcheng, et al.
Published: (2026)
by: Xu, Pengcheng, et al.
Published: (2026)
MME-Unify: A Comprehensive Benchmark for Unified Multimodal Understanding and Generation Models
by: Xie, Wulin, et al.
Published: (2025)
by: Xie, Wulin, et al.
Published: (2025)
MedGround-R1: Advancing Medical Image Grounding via Spatial-Semantic Rewarded Group Relative Policy Optimization
by: Xu, Huihui, et al.
Published: (2025)
by: Xu, Huihui, et al.
Published: (2025)
From Summary to Action: Enhancing Large Language Models for Complex Tasks with Open World APIs
by: Liu, Yulong, et al.
Published: (2024)
by: Liu, Yulong, et al.
Published: (2024)
How Long Can Unified Multimodal Models Generate Images Reliably? Taming Long-Horizon Interleaved Image Generation via Context Curation
by: Chen, Haoyu, et al.
Published: (2026)
by: Chen, Haoyu, et al.
Published: (2026)
Rethinking Token-Level Policy Optimization for Multimodal Chain-of-Thought
by: Li, Yunheng, et al.
Published: (2026)
by: Li, Yunheng, et al.
Published: (2026)
Contextual AD Narration with Interleaved Multimodal Sequence
by: Wang, Hanlin, et al.
Published: (2024)
by: Wang, Hanlin, et al.
Published: (2024)
LaneCorrect: Self-supervised Lane Detection
by: Nie, Ming, et al.
Published: (2024)
by: Nie, Ming, et al.
Published: (2024)
Towards Harmless Multimodal Assistants with Blind Preference Optimization
by: Li, Yongqi, et al.
Published: (2025)
by: Li, Yongqi, et al.
Published: (2025)
HiRes-LLaVA: Restoring Fragmentation Input in High-Resolution Large Vision-Language Models
by: Huang, Runhui, et al.
Published: (2024)
by: Huang, Runhui, et al.
Published: (2024)
Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning
by: Zhang, Lei, et al.
Published: (2026)
by: Zhang, Lei, et al.
Published: (2026)
iFlame: Interleaving Full and Linear Attention for Efficient Mesh Generation
by: Wang, Hanxiao, et al.
Published: (2025)
by: Wang, Hanxiao, et al.
Published: (2025)
OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text
by: Li, Qingyun, et al.
Published: (2024)
by: Li, Qingyun, et al.
Published: (2024)
VINO: A Unified Visual Generator with Interleaved OmniModal Context
by: Chen, Junyi, et al.
Published: (2026)
by: Chen, Junyi, et al.
Published: (2026)
MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer
by: Tian, Changyao, et al.
Published: (2024)
by: Tian, Changyao, et al.
Published: (2024)
M$^{2}$Chat: Empowering VLM for Multimodal LLM Interleaved Text-Image Generation
by: Chi, Xiaowei, et al.
Published: (2023)
by: Chi, Xiaowei, et al.
Published: (2023)
Region-Constrained Group Relative Policy Optimization for Flow-Based Image Editing
by: Ouyang, Zhuohan, et al.
Published: (2026)
by: Ouyang, Zhuohan, et al.
Published: (2026)
Learning to Generate via Understanding: Understanding-Driven Intrinsic Rewarding for Unified Multimodal Models
by: Pan, Jiadong, et al.
Published: (2026)
by: Pan, Jiadong, et al.
Published: (2026)
DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies
by: Song, Wei, et al.
Published: (2025)
by: Song, Wei, et al.
Published: (2025)
Group Critical-token Policy Optimization for Autoregressive Image Generation
by: Zhang, Guohui, et al.
Published: (2025)
by: Zhang, Guohui, et al.
Published: (2025)
Query-Kontext: An Unified Multimodal Model for Image Generation and Editing
by: Song, Yuxin, et al.
Published: (2025)
by: Song, Yuxin, et al.
Published: (2025)
Similar Items
-
Reason2Drive: Towards Interpretable and Chain-based Reasoning for Autonomous Driving
by: Nie, Ming, et al.
Published: (2023) -
KFFocus: Highlighting Keyframes for Enhanced Video Understanding
by: Nie, Ming, et al.
Published: (2025) -
SlowFocus: Enhancing Fine-grained Temporal Understanding in Video LLM
by: Nie, Ming, et al.
Published: (2026) -
SemHiTok: A Unified Image Tokenizer via Semantic-Guided Hierarchical Codebook for Multimodal Understanding and Generation
by: Chen, Zisheng, et al.
Published: (2025) -
UNIT: Unifying Image and Text Recognition in One Vision Encoder
by: Zhu, Yi, et al.
Published: (2024)